A document arrives from a colleague. It is well structured, the prose is clean, the argument moves. Somewhere inside it is a claim you would be embarrassed to repeat in front of a client if it turned out to be wrong.
So you ask yourself a question that did not need asking three years ago. How was this made?
You cannot tell. You do not know what was asked of the model, what came back unedited, or what your colleague verified and what they took on faith. Which leaves two options. Verify it from the beginning, reconstructing work that has already been done. Or accept it on trust and carry whatever it contains forward under your own name.
That is where the productivity goes.
Your colleague was genuinely faster. They are not exaggerating and the time they saved was real. It has simply been transferred to you, as verification you cannot skip and cannot bill for. Multiply that across a workflow of six people and the organisation's net gain rounds to zero while everyone in the chain honestly reports being quicker. Everybody is telling the truth. The organisation still has nothing.
I suspect that is a large part of what the failure statistics are actually measuring. The binding constraint is not model capability, and it is not implementation methodology. It is whether one person can trust another person's AI-assisted work enough to build on it without redoing it.
What makes it stubborn is that the fix has the shape of a commons problem. Describing how you made something costs you time and delivers the benefit to somebody else's desk. In an organisation that measures individual output, which is to say in most organisations, every incentive points towards producing more and explaining less. The rigorous document and the fluent one look identical in the metrics, and the fluent one arrived first. Until the measurement can tell them apart, rational individuals will keep eroding the trust the team runs on, and no purchase order will touch it. It is a question of what gets rewarded, which is to say a management problem, and exactly the kind the current enthusiasm for technology keeps stepping around.
A tradesman who handed over work in this condition would not stay in business. The morass of half-finished thinking passed between knowledge workers under the banner of AI-assisted productivity would shame a plumber.
I have seen the solved state, but only at small scale, and the reason it works there is instructive. In a team of four who share a room, verification stays cheap: the context is shared, the work is visible, everyone can check everyone. The practices holding it together are unglamorous. A note travels with the work saying what was asked of the tool and what a person checked. Model-generated sections are marked as such, with a line on what was verified. Nobody calls it governance; it is just what you do so a colleague can extend your work instead of rebuilding it.
Scale breaks it because verification cost grows faster than headcount. Every handover adds another point where somebody must establish quality without having been present when the work was made.
What I cannot hand you is the mechanism for doing this across teams that do not share a room, to standards that do not yet exist, under incentives that currently reward volume over verifiability. That is a design problem in its own right and it is still open on my desk. But naming it precisely is the precondition for solving it, and most organisations have not named it at all. They are still buying licences and wondering why the individual gains never add up to anything.
When AI-assisted work does transfer with real confidence, value stops leaking at the seams and starts to compound. Each contribution builds on the last rather than replacing it, and the output reflects the expertise of the team rather than the enthusiasm of its fastest member. That compounding is the flywheel. Without it, you have a collection of personal productivity gains and a very expensive year.
Something to try this week. Pick one team and add three lines to the top of every AI-assisted document for a month. What you asked the tool to do. What you checked yourself, and against what. And what you would not yet stake your name on.
The third line is the one people find hard to write. That difficulty is the finding.