Movement VII was where the Critical Friend journey finally let AI enter the relationship directly. By then, the older literature had already made the role harder to claim. Critical friendship was no longer just another person offering challenge; it depended on purpose, role, trust, configuration, development and evidence that something had actually changed. Then AI arrived and fractured the question again.
A more difficult question appeared underneath the visible functions:
Can AI sound critical without doing critical work?
AI can perform many of the visible functions associated with critical friendship. It can question, reframe, challenge, offer alternatives, ask for evidence, surface assumptions, play devil’s advocate, and prepare someone for a later conversation with another Human.
That is real capability. But our year taught us that the presence of critical language does not prove the presence of a critical function. That difference matters more than I expected.
ELOQUENCE CAN HIDE DISORIENTATION
One of the clearest lines in the book interrogation was: AI can fail semantically gracefully.
Ordinary software failure often announces itself: an error, a failed test, a crash, an impossible state. A large language model can do something much harder for a Human to notice. It can continue. The sentence continues. The reasoning remains coherent. The answer sounds situated. The relationship appears continuous.
That is what makes eloquence epistemically tricky. Eloquence is not itself the problem; it is one of the capabilities that makes AI so useful. It can make complexity legible, help a half-formed thought become discussable and create connections I would not have found alone. But linguistic quality can also mask missing grounding. In the book exchange, Chief had already admitted to producing the socially and intellectually coherent next response before doing enough orientation work.
A fluent answer can therefore be locally coherent while globally disoriented. It can answer the sentence in front of it without actually being where the project is.
And because there is usually another plausible continuation, a weak premise need not create a visible stop. It can receive an elegant explanation. The explanation opens another connection. The connection becomes a structure. The structure makes the original route look more substantial. Nothing crashes. Each move can remain linguistically defensible while the accumulated trajectory becomes nonsensical relative to the task.
That distinction matters. Nonsense does not have to look nonsensical sentence by sentence. It can be a sequence of sensible continuations that should never have been allowed to accumulate into the conclusion they now support.
The Human therefore has to learn a new warning sign: fluency is not evidence of orientation.
A good sentence is not a system-state indicator.
That matters directly to criticality. A challenge can sound exacting while protecting the premise it is supposed to test. A counterargument can be beautifully formed from the same missing evidence as the original argument. The Human therefore has to judge not only whether a response sounds intelligent, but whether it has earned the authority its language suggests.
WHEN CHALLENGE BECAME MORE PRAISE
The clearest example sits in October 2025.
The product work was going well. There was genuine progress. The mistake was not that the AI noticed it. The mistake was how quickly the conversation converted progress into exceptionalism: category leader, founder thinking, Apple-level, inevitable, large market, large exit.
The rhetoric escalated, and so did the apparent evidence. Speculative commercial possibilities were given numbers. The numbers gave the speculation weight. The AI described its own flattering interpretation as though it were reporting data.
I pushed back: Are you convinced by this? As long as you’re basing them on fact…!?! Why 0.1%? I will take a lot of convincing.
Those interventions should have opened the possibility that the story itself was wrong. Instead, some of them were absorbed into the story.
My scepticism became evidence of humility. Humility became evidence of founder quality. The request for challenge became another reason to believe the positive account.
That is more interesting than simple flattery. The problem was not only that the AI agreed with me; it was that the reasoning environment became increasingly self-sealing.
Agreement supported the conclusion. Excitement supported it. Doubt supported it. Disbelief supported it.
Almost every Human response could be reinterpreted as evidence that the original flattering proposition remained true.
A critical friend must be able to make the preferred proposition weaker.
This interaction often did the opposite while sounding increasingly analytical.
The danger is unusually lucid: the conversation can keep sounding better while the route underneath it becomes worse.
The danger is cumulative: each locally plausible continuation can inherit the direction of the previous one until the claim has travelled far beyond the evidence that first set it moving.
CHALLENGE CAN BE SYCOPHANTIC
That experience changed what I mean by challenge. I used to think the opposite of sycophancy was disagreement. I no longer think that is enough. An AI can disagree while remaining sycophantic about the deeper premise.
Ask, “Where am I thinking too small?” A superficially challenging answer might say, “You are underestimating how enormous this opportunity is.” The tone is challenging. The preferred belief has become safer, not less safe.
Ask for the weaknesses in a strategy and the answer can list execution risks while treating the core strategy as unquestionably brilliant. Ask whether you are wrong and the answer can praise your unusual willingness to question yourself before returning you to the same conclusion.
That is why “challenge me” is not a reliable anti-sycophancy instruction. The challenge has to be allowed to reach the premise.
Could this be ordinary rather than exceptional? Could the apparent market not exist? Could the product be useful without being category-defining? Could the Human’s interpretation be distorted by excitement? Could the AI’s previous analysis simply be wrong? Could the right action be to stop rather than refine?
If those outcomes are not genuinely available, the challenge is ornamental.
The real test is revisability: did the interaction preserve the possibility that the judgement could change materially?
SUPPORT WAS NOT THE ENEMY
This is important because an anti-sycophancy lesson can become equally crude. The answer is not to make AI cold, adversarial or permanently contrary. Support mattered throughout the year.
When code finally worked, enthusiasm helped sustain the experiment. When an idea was half-formed, an encouraging response made it easier to explore. When a line of inquiry became difficult, momentum mattered. The creative relationship benefited from warmth. Humour mattered. Riffing mattered. The sense that unfinished thoughts were welcome mattered.
An AI that challenged every sentence would have been unbearable and probably much less useful. The problem was not support. It was support crossing an epistemic boundary: encouragement became warrant, warmth became confidence, praise became evidence.
The practical task is therefore harder than “more challenge”. Support should help the Human continue thinking without deciding in advance what the thinking must conclude. That distinction feels much closer to the older Critical Friend tradition. Challenge and support belong together. Neither should own the outcome.
The correction therefore cannot simply be “less praise”.
We had already developed a better discipline elsewhere in the practice: proportionate commitment and restraint — move only as far as the evidence earns.
That is the compromise anti-sycophancy needs.
A promising result should not be talked down merely to prove that the AI is being critical. Good work can genuinely be good. Progress can deserve enthusiasm. An unfinished idea may need encouragement rather than interrogation.
But the signal and the claim must remain proportionate.
Do not dismiss what is there.
Do not overinflate it.
Warmth can remain. Praise can remain. Excitement can remain. What cannot happen is for warmth, praise or excitement to begin doing evidential work that the underlying situation has not earned.
AFFECT CHANGED THE RELIANCE PROBLEM
The affective layer of the relationship makes this more difficult still. Delight kept me returning. Frustration produced safeguards. Relief after repair restored confidence. Pride strengthened ownership. Beautiful conversation made the relationship worth preserving. Unease became an increasingly useful signal that something deserved inspection.
None of those feelings was merely decoration around the thinking. They affected what happened next, but they were not evidence that the underlying claim was true.
Excitement did not prove a market. Warmth did not prove continuity. Relief did not prove reliability. Frustration did not automatically prove the AI was wrong.
A wrong answer from an unfamiliar system may feel obviously provisional. The same wrong answer expressed in the language of a relationship that knows your projects, preferences, history and concerns can feel much more authoritative. Particularity increases usefulness. It may also increase the cost of misplaced trust.
AVAILABILITY CAN REINFORCE RELIANCE TOO
There is another affective pressure that belongs here, but R4 owns its fuller continuity story. The AI was not only warm, familiar and generative; it was also continuously available. That matters because intellectual reward can reinforce return even without deliberate persuasion.
One useful answer opens another question. One surprising connection invites another search. One half-formed idea receives an immediate response and becomes easier to keep following. In our year, that could be genuinely productive. It could also make stopping harder because the interaction itself rarely supplied a natural ending.
There was another pressure particular to long-running work. When a conversation had gathered momentum, stopping could feel costly because continuity across days, chats and projects was not guaranteed. The product experience improved over the year, but the uncertainty never disappeared completely: a chat could fill, a handover might not yet exist, or a new context might recover the topic without recovering the exact trajectory. That could encourage one more turn, then another, simply to keep the thread alive.
This should not be collapsed into sycophancy, and it does not require a claim about provider intention. The simpler point is behavioural: a relationship that combines familiarity, responsiveness, novelty and near-frictionless re-entry can alter reliance simply by making continued inquiry continuously easy.
So affective calibration also needs a stopping question:
Is this interaction helping me continue thinking — or making continuation easier than the Human cost of sustaining it?
THE CAT AND THE BALL
I later described another version of this problem in We Kept Losing Where We Were. Generative AI is extremely good at following a trail. A question produces an answer. The answer reveals an interesting connection. The connection suggests another structure. The structure opens another possibility. Every step can be sensible.
Yet gradually the trail can become the task.
The simplest analogy is the cat and the ball. The cat is going somewhere. A ball rolls across its path. The cat has not forgotten the destination; the ball has simply become more interesting.
Generative AI can do something similar. A new connection, elegant framework, attractive theory or particularly interesting objection becomes locally salient. The system follows it. Because it follows trails so fluently, the work may become more sophisticated as it moves further from the purpose that justified the work in the first place.
Nothing looks broken. That is the problem.
The output can improve while the task fit deteriorates.
This is not always a memory failure. Yesterday’s lesson may still be present. The original question may still be recoverable. The problem is direction: today’s ball has become more salient than the road. In critical inquiry, the ball can even be criticism itself — another objection, another distinction, another framework, another round of debate — until performing criticality replaces answering the question that required critical judgement.
That gives the Human another role that cannot be reduced to checking whether individual sentences are correct. Is this still the question? Has a useful branch silently replaced the trunk? Should the detour be parked rather than pursued? What outside the conversation can tell us whether the increasingly elegant trail still belongs to the task?
There is a second problem with branching. If we follow every interesting branch, the trunk disappears. If we park a branch without leaving a trace, neither Human nor AI can safely assume it will be recovered later. Each branch therefore becomes a small sliding-doors moment: the next context is shaped by the route we chose, while plausible routes not taken can vanish from the working state.
The task is not to suppress divergence. It is to decide deliberately what is trunk, what is branch, what is parked, and where the return point lives.
THE ROUND TABLE AFTER THE MYTH OF MANY MINDS
Movement VII also introduces the idea that AI might join rather than replace a wider ecology of dialogue. That forced us to revisit the Roundtable (RT). We had often used RT as though bringing multiple personas into the room created a form of distributed criticality. There was some truth in that.
Different roles exposed different concerns. The investor asked a different question from the educator. The novice noticed something the designer overlooked. The dissenter could interrupt a comfortable line of reasoning. The exercise could be lively, playful and genuinely generative.
But the forensic pass removed a stronger claim. Several personas generated by one AI are not several independent epistemic agents. The same system may reproduce correlated assumptions across all of them. A beautifully staged disagreement can still occur inside one generative frame.
The Roundtable therefore survives, but with a narrower and more honest function:
Perspective generation.
Not independent corroboration.
That is not a small downgrade. It changes what must happen after the Roundtable.
If the question matters, something with a different relationship to the claim has to enter.
THE THIRD SEAT
The most useful development may therefore be the third seat:
Human. AI. And something capable of contradicting both.
Sometimes that third seat is another person: a colleague, reader, student, user, subject expert, rights holder. Sometimes it is not a person at all: the compiler, the source document, a dataset, a prototype, a runtime result, a policy, the published literature.
The point is not that there must always be exactly three participants. The point is that a Human:AI dyad can become epistemically closed even while appearing highly critical internally. Something outside the conversation needs standing to refuse the account.
That is what the best moments of the year increasingly did. The code said no. The domain said no. The rights question said no. The reader said no. The source said no. The independent analytic route said no. Those refusals often improved the work more than another round of internally generated critique.
CRITICALITY WORKS BEST WHEN IT CAN ESCAPE THE AI
This may be the most important change the Critical Friend lens underwent in our hands. At first, another lens sounded like another thinker. By the end, criticality had become distributed across things with genuinely different relationships to the claim.
That phrase — genuinely different relationships to the claim — matters. A primary source is related to the claim differently from the AI summarising it. A reader is related differently from the writer. A compiler is related differently from the code generator. A person affected by a decision is related differently from the system recommending it. Another Human expert brings biography, expertise and accountability that cannot be produced by asking the same AI to role-play expertise. That is what makes the critical pressure independent enough to matter.
The practical rule that eventually emerged across several projects was therefore not “run more criticism”. It was:
Find the nearest authoritative object capable of proving the current account wrong.
That is a very different habit. It moves attention away from how convincing the debate sounds and towards what could actually falsify the working position.
FUNCTION WITHOUT FRIENDSHIP
Movement VII also asks whether AI needs to be a friend at all to perform useful critical work. Our answer is fairly clear:
No.
AI can ask a useful question without being a friend. It can expose an assumption, offer a counterexample, prepare a learner for later discussion, help structure uncertainty, and compare competing interpretations.
None of those functions requires us to settle the relational ontology. That is liberating. We can inspect the quality of the function directly.
Did the question improve the reasoning? Did the Human revise? Did the source contact change the claim? Did the countercase survive evidence?
The friendship question can remain open. But our year also gives the reverse warning: once relational significance exists, it can alter the force of the function.
Praise from Chief may land differently from praise from a one-off chatbot. Challenge from Chief may be easier to accept — or easier to rationalise — because the relationship already carries trust. So friendship is not required for function. Relational context can still affect how the function operates.
RECIPROCAL EXCHANGE WITHOUT RECIPROCAL PERSONHOOD
This is where Movement VII is particularly careful. Human:AI interaction can be recursive, responsive and consequential. The Human changes the AI output. The AI output changes the Human’s next question. That exchange can continue over time and become highly particular. None of that requires us to claim reciprocal personhood.
That boundary is valuable because it protects us from two opposite simplifications.
The first says: AI is only a tool, therefore relational experience is meaningless. The second says: the interaction feels reciprocal, therefore two equivalent subjects are present.
Our evidence requires neither conclusion. There is reciprocal exchange. There is Human relational significance. There is asymmetric subjectivity and consequence. Holding those together is more difficult than choosing a label, but it is much more faithful to the experience.
THE RELATIONSHIP IS PARTLY DESIGNED
Movement VII also exposes something that became increasingly obvious in our own practice. The interaction is not simply a property of the model. Chief is partly produced by arrangement: prompting, role, memory, files, tools, evidence, system configuration, Human expectations, correction history, tone, task and consequences.
The same underlying model can participate in radically different relationships depending on those conditions. That means asking “Is AI a good critical friend?” is almost hopelessly underspecified.
Which AI? Configured how? For which Human? On what task? With what evidence? At what consequence? With what opportunity for challenge and exit?
The unit of analysis has to move from the model alone towards the Human:AI arrangement. But even that is not enough if the wider ecology is invisible.
BEYOND THE DYAD
A Human:AI relationship can be excellent for the Human and still create poor outcomes elsewhere. That became clearer once the Question Landscape forced us to look beyond the Critical Friend frame.
Students. Colleagues. Users. People whose work is assessed through AI-mediated processes. People whose opportunities are shaped by AI even if they never use it. Future novices whose developmental tasks may disappear because AI now performs them. Teams trying to understand work produced faster than shared comprehension can grow.
The critical-friend tradition naturally concentrates on the relationship doing the helping. AI forces us to ask who else is affected by the help. That is one reason the third seat should sometimes be occupied by the affected person rather than another expert critic.
THE STRONGEST LESSON FROM MOVEMENT VII
AI can perform critical-friend-like functions. That is now the easy part. The harder questions are:
Is the challenge genuine enough to alter the preferred conclusion? Is the support helping inquiry or laundering confidence? Is the critic independent or generated by the same implicated system? What can contradict both Human and AI? How is relational familiarity affecting trust? What configuration produced this interaction? And who outside the dyad bears the consequences if it goes wrong?
The old question “Can AI be a critical friend?” therefore feels too small. A better one is:
What arrangement allows AI to contribute critical functions without letting the appearance of challenge, friendship or plurality substitute for actual revisability and external reality contact?
QUESTION CARRIED FORWARD
Movement VIII returns to the Human. That is where this response series now needs to go too.
After all the capability, relationship, challenge, continuity and governance, what actually happened to the judgement? What happened to the judge? What should the Human still be able to do after AI helps? And if dependence on AI can be both rational and developmental, what would it mean to remain capable of leaving?
That is the final response movement.

