How do we know whether a Human:AI arrangement actually improved the judgement?
Not the document.
Not the fluency.
Not the speed.
The judgement.
That distinction has been sitting underneath this project from the beginning.
AI makes visible work easy to admire.
A cleaner report.
A better structure.
More evidence gathered.
More alternatives considered.
A persuasive explanation.
Those improvements can be completely real.
But a better artefact does not automatically tell us whether the decision behind it became better warranted.
That was the disturbance behind The Neighbourhood.
The finished work had stopped being a reliable proxy for everything we wanted to know about the person and the process that produced it.
The critical-friend journey sharpens the same problem from another direction.
A helpful interlocutor is not valuable merely because they contributed more words.
They are valuable when the inquiry changes in a way that survives scrutiny.
So what would count as improved judgement?
Sometimes it is simple.
An error is found.
A missing source changes the conclusion.
A counterexample causes a recommendation to be narrowed.
A calculation is corrected.
A better option appears.
The world later confirms that the decision worked.
But many judgements do not offer that kind of immediate scoreboard.
Then improvement has to be inspected more carefully.
Our working practice began separating things that fluent AI tends to blur:
evidence
interpretation
inference
recommendation
judgement
That separation matters because each can fail differently.
The source can be wrong.
The interpretation can overreach the source.
The inference can travel too far.
The recommendation can ignore consequences.
The final judgement can be perfectly coherent and still belong to the wrong person.
AI can contribute to every stage.
It should not make the stages disappear.
A second test is whether the inquiry encountered resistance.
Did the AI merely elaborate the starting position?
Or did something have a genuine chance to change it?
A dissenter.
A negative case.
A source that did not fit.
A reality check.
A person affected by the decision.
The strongest Human:AI work in our own record often did not improve because the AI supplied a final answer.
It improved because a provisional answer was rejected, bounded, reopened or sent back to the evidence.
That is why disagreement matters.
Not because friction is virtuous.
Because a judgement that has never met anything capable of changing it has not been tested very hard.
Then comes authority.
An AI-generated recommendation can be excellent.
Who decides whether to act on it?
If the answer is “the human”, that statement only means something if the human can understand enough of the state of play to intervene.
Ceremonial approval is not judgement.
A person pressing accept on a conclusion they cannot inspect has retained a button, not necessarily authority.
This is where evidence from AI-supported learning becomes useful.
Gao and Zhang’s 2026 study of doctoral researchers distinguishes coherent AI-supported output from defensible understanding.
Learner-controlled engagement was more visible when students continued to read, verify, reconstruct and adapt the generated material until they could explain and defend the reasoning behind it. Their study does not give us a universal test for judgement, but the distinction travels well.
Can the person reconstruct why this conclusion is warranted?
Can they explain what evidence would change it?
Can they identify what came from AI?
Can they say where uncertainty remains?
Can they defend the decision to somebody who disagrees?
Those questions inspect more than the surface of the artefact.
There is also reality contact.
A Human:AI conversation can become extraordinarily coherent while remaining wrong about the world.
Sources matter.
Data matter.
People matter.
Implementation matters.
What happens after the recommendation matters.
The loop should sometimes break outward.
Then what comes back from reality can change the next Human:AI cycle.
That is not a failure of the thinking relationship.
It is part of what keeps it honest.
So the evaluation of judgement begins to look less like:
Was AI used?
or:
Was the answer good?
and more like:
- What decision was actually being made?
- What evidence supported it?
- What alternatives could have changed it?
- What uncertainty remained?
- Who had authority?
- Could that person interrupt or refuse?
- What happened when the judgement met reality?
None of this requires the human to do every part of the work.
AI may be the reason the judgement improved.
It may retrieve the source nobody remembered.
Expose the countercase.
Run the comparison.
Hold a larger evidence field than one person could manage alone.
The point is not to preserve human labour.
It is to preserve the conditions under which a consequential judgement can still be owned, challenged and revised.
And even that is only half of the evaluation.
Because a Human:AI arrangement can improve today’s judgement while changing tomorrow’s judge.
Even if the judgement improved, what happened to the person making it?
References
Gao, L. and Zhang, F. (2026) ‘Cognitive offloading and metacognitive calibration in generative AI-mediated doctoral learning: a grounded theory study of accountable reconstruction’, Frontiers in Psychology, 17, 1920947. doi:10.3389/fpsyg.2026.1920947. Open-access article.

