Three findings from the past eighteen months, all measured rather than asserted.
Roughly half of newly published web articles are now primarily AI-generated. The share crossed fifty percent in early 2025 and has sat there since.
Forty percent of employees report receiving low-substance AI-generated work from a colleague in the previous month. Researchers gave it a name: workslop. Each incident consumed close to two hours to untangle. Self-reported, so hold the precision loosely, though the direction is not subtle.
And the MIT NANDA study putting 95 percent of generative AI pilots at no measurable financial impact. That one is preliminary, unreviewed, and its methodology has been publicly contested. I use it as one signal among several rather than a headline.
The three findings share one feature. In no case did the machine produce nothing; in every case it produced something that looked right.
That distinction is the whole problem, and it is narrower than the usual complaint. The usual complaint is that these systems are unreliable, which invites the reply that they are getting better every month. Probabilistic behaviour is not the difficulty. Plenty of systems are probabilistic and reliable. The difficulty is that the output does not signal its own error. The failure mode is not an error message. It is a well-formed paragraph that happens to be wrong.
In almost every other tool an organisation depends on, failure announces itself. The build breaks, the machine stops, the number is visibly wrong. Confidence and correctness travel together closely enough that a competent person can triage on surface signal alone, which is exactly why organisations have run on review-and-approve workflows for a century and why those workflows worked.
Generative AI severs that link, quietly and completely. Most quality processes have not registered it, because for the entire working lives of the people who designed them, fluency was evidence of competence.
Which is why the standard response fails. Workslop does not look like slop. It looks like work. Conscientious people wave it through, and telling them to review more carefully has no purchase whatsoever, because carefulness was never the missing ingredient. The signal was.
It is worth being honest about where the opposite expectation came from. The magic button, one prompt in and a finished corporate asset out, was sold. The most confident claims about the imminent end of knowledge work have been made by people holding the largest positions in the outcome, into a capital-raising environment, and they should be read accordingly. The technology is remarkable. It was never the oracle in the pitch.
Meanwhile the volume climbs. When the marginal cost of generating a document approaches zero, organisations do not produce the same number of documents more cheaply. They produce vastly more documents. Generation has become cheap and judgment has not, so the work of sorting the good from the merely plausible falls entirely on human attention, which is the one input that has not become cheaper at all.
The bottleneck has moved. It is no longer whether we can produce the thing. It is who can tell whether it is any good.
Something to try this week. Take one AI-assisted document that is about to go out and give yourself twenty minutes with a different instruction. Not review it. Assume there is something wrong in it, and go and find that thing.
It is a different activity and it feels different within a minute. Reviewing has no stopping rule, so it ends when you run out of patience. Searching ends when you find something, or when the clock does.
In my own work what I find is almost never what I would have caught by skimming, and it usually turns up in the first ten minutes. That is practice rather than evidence. It costs twenty minutes to test on your own desk.