I have been saying something for about two years that is not quite right, and since this series is partly about verification, I should correct it in public.
The claim was that the quality of AI output is capped by the critical-thinking capacity of the operator. It sounds right and it is imprecise. AI can plainly produce work exceeding what a given person could produce unaided, which is most of why anyone uses it. If quality were capped by the operator, that could not happen.
What is capped is the operator's ability to recognise quality. You cannot accept, reject, improve or safely deploy what you cannot evaluate. The ceiling is not on generation. It is on verification.
That matters more than a corrected phrase, because it collides with an assumption nobody examines. In most domains, verifying is cheaper than producing. It is easier to check a proof than to find one, easier to taste the soup than to cook it. That asymmetry sits underneath every human-in-the-loop diagram ever drawn, and it is why review-and-approve has been a sensible division of labour for as long as organisations have existed.
For fluent-and-wrong output, the asymmetry inverts. When plausibility is high and correctness is unknown, genuine verification means reconstructing enough of the underlying reasoning to establish whether the answer is right. That is most of the work of producing it, plus the burden of resisting a text that is actively signalling its own competence. Checking becomes more expensive than doing.
I should be accurate about where this comes from. When I first formulated it from practice I believed it was mine. It is not. Over the past eighteen months, research from several directions has reached the same inversion: software engineering, where review effort now measurably exceeds production effort; scientific workflows, where generation costs collapsed and verification costs did not; and the economics of AI, in the distinction between checkable outcomes and context-sensitive judgment. I offer it as synthesis, and claim only what follows from it operationally.
The first thing it explains is the reviewer trap, and better than I managed before. The usual account is psychological: attention drifts, fluent text invites lazy reading. Both are true. But the deeper reason the arrangement collapses is economic. You handed the human the more expensive job, in less time, with worse information than the machine had, and then measured them on throughput. Of course it decays into approval. It was never affordable as designed.
Now the evidence that cuts against me, which is the most interesting part of this. Two of the most-cited studies in the field found the opposite of what I would predict. In a field experiment with consultants, the largest gains went to below-average performers, compressing the distribution; a study of customer support agents found the same shape. If output were bounded by the operator's judgment, AI should widen the gap between strong and weak performers rather than close it.
The reconciliation runs along the same axis. Both examine work where the quality of an answer is easy to establish: a support issue is resolved or it is not, and a graded task has a right answer. Where output is readily verifiable, AI compresses performance, because checking is cheap for everybody and the tool supplies what the weaker operator lacked. Where output is hard to verify, the operator's judgment is the only available check, and there I expect the distribution to widen. Worth noting that in the consulting experiment, tasks deliberately placed outside the model's competence saw assisted consultants perform worse than unassisted ones.
The two halves of that claim do not have equal support. The compression half is well evidenced. The widening half is my inference, consistent with the published work but not demonstrated by it. I would rather it were tested than believed.
Something to try this week. List the five things you are currently using AI for, or planning to. For each, ask one question with a clock attached: could a competent colleague tell within ten minutes whether this output is right?
Drafting against a defined template usually passes. Cross-document consistency checking usually does not. Anything you would let the machine pass or fail on its own authority comes off the list altogether.
Then look at where the failures cluster. The uncomfortable pattern, in my experience, is that the tasks where AI most flatters your weakest performers are the ones to worry about least, and the trouble gathers exactly where nobody can easily tell how well they are doing.