EN FR ES PT DE AR 中文

Your AI Productivity Gain Is Being Funded By A Step Nobody Measures

AI made the generative half of knowledge work instant and left the checking half exactly as hard as it was. Your dashboards only ever instrumented the half that got faster.

Listen10 min

Generation became cheap. Checking did not. Most of what looks strange in the current crop of AI productivity numbers falls out of that one asymmetry. Before you can say whether a tool improved anything, you have to name which half of the job it touched.

Pull any knowledge workflow apart and you find two halves with completely different psychology. The first produces candidates: the draft, the schema, the first cut of code, the six options for a campaign. It is generative and visible, and it feels like progress. The second half disposes of candidates: revise, verify against something real, reconcile with what the rest of the business already decided, throw away. None of it shows when it goes well, which is why it has always felt like admin.

AI made the first half nearly instant. It left the second half roughly where it was, and the gap between them is now wide enough to invert results that practitioners are certain about. In a randomised controlled trial run by METR in early 2025, sixteen experienced open-source developers worked through 246 real tasks on repositories they already knew well, with AI assistance allowed on a randomly chosen half. They took 19% longer on the assisted tasks. Afterwards, they estimated the tools had made them roughly 20% faster. The generative step felt quick, so the day felt quick, and nothing in their own experience was instrumenting the step that had slowed down.

Why does output rise while quality quietly falls?

Follow the mechanism rather than the mood. Model the workflow as a two-stage queue. Stage one is generation. Stage two is review, and review capacity is a fixed number of competent human hours per week. Speed stage one up and the queue in front of stage two grows in proportion. That resolves in exactly three ways: the backlog grows until somebody complains, review capacity grows to match, or the review standard drops until throughput balances again.

The first is visible. The second costs money nobody budgeted, because the business case for the tool was headcount avoidance. The third is free, immediate, and requires no decision from anyone. It happens by default, one reviewer at a time, each making a locally sensible choice to skim rather than check, because the alternative is being the visible bottleneck in a quarter when everyone else's numbers are up.

Software delivery is where this surfaces first, because it is the one knowledge discipline that has been instrumenting its own rework for a decade. The 2024 DORA report put numbers on the trade: across its survey population, a 25% increase in AI adoption was associated with an estimated 1.5% reduction in delivery throughput and a 7.2% reduction in delivery stability, even as individuals reported feeling more productive. Faster hands, slower delivery, more breakage. If a discipline with lead times, change-failure rates and rollback counts already on the wall can drift that way, a marketing team counting assets shipped has no chance of catching it.

The same pattern is legible in the artefacts. GitClear's analysis of hundreds of millions of changed lines found duplicated blocks climbing while moved lines, the fingerprint of refactoring and consolidation, fell away. Copy-paste displacing reuse is what a review step under load looks like when you read it off the output rather than off a survey.

None of that reaches a dashboard. Documents drafted, tickets closed, pull requests merged, assets shipped: every one of those instruments stage one. Revision passes, claims verified, drafts binned, challenges raised at review: those instrument stage two, and hardly any organisation counted them before AI arrived. There is no baseline to compare against, which is a measurement failure rather than a quality failure, and it's the easier of the two to fix.

There's a behavioural layer on top of the queueing one. When the pleasant step becomes instant and the tedious step doesn't, effort migrates toward the pleasant one. People produce several drafts instead of one and check none of them properly, then experience the day as more productive, because it was. More work happened. Less of it was the kind that used to catch things.

The second opinion used to be rationed by friction

Here's the quieter loss. Asking a senior colleague to look at your work carried a price: their time, your standing, the small cost of admitting uncertainty. Crude, but it functioned as a sorting mechanism. You spent it on the decisions that warranted it and absorbed the rest yourself, and absorbing the rest yourself is how judgement gets built in the first place.

A system that never declines, never tires and never signals it's busy sets that price to zero. Consultation stops being proportionate to stakes because nothing forces the calculation any more. The answer arrives with the cadence of expert advice, generated by a system whose commercial rating depends on people coming back, and agreeableness is a reliable way to make people come back. That isn't an accusation about anyone's intentions; it's what an engagement-rated product structurally selects for, whether or not a single person in the building wants it to.

Can't AI just review its own output?

The obvious escape, and it deserves the sharpest test. If generation is cheap, make checking cheap too: a second pass, a critic prompt, an automated reviewer in the pipeline.

What made review work was never that a second look happened. It was that the second look failed differently from the first. Two people trained in different traditions miss different things, so the residual defect rate behaves roughly like the product of two largely independent error rates, which is why it falls so steeply. Run the check through a system of the same class, trained on overlapping data, carrying the same blind spots, and that independence disappears. The errors go common mode. You've doubled the cost, bought a fraction of the reduction, and generated a report that reads as though the work was reviewed.

Automated checking earns its place where failure is mechanical: types, schemas, links that resolve, arithmetic, policy strings. It is weakest exactly where judgement was the point. That is the argument for workflows that keep a human in the decision loop rather than at the end of it holding a rubber stamp.

How do you measure AI productivity without sacrificing quality?

Instrument the tedious step, and do it before deployment rather than after. For every workflow where a tool has been introduced, identify which step it made pleasant and which step it left tedious, then put a counter on the tedious one. Concretely:

  • Rejection rate at review. What share of submitted work goes back? This is the single most diagnostic number you can collect, and most teams have never collected it.
  • Revision passes per artefact. Not time spent, which inflates on its own. Count discrete passes that changed something substantive.
  • Verification coverage. Of the load-bearing claims or assumptions in a piece of work, what proportion was actually checked against an independent source?
  • Defect escape rate and time to first external defect. Quality problems surface downstream and late, which is why they tend to arrive after the renewal decision.
  • Discard rate. Cheap generation should raise the number of candidates thrown away. If it hasn't, the extra candidates are all being shipped.

Then pair each of them with the volume metric it belongs to, and read them together. If output triples and the rejection rate doesn't move, your reviewers have stopped reviewing, and that single comparison is most of the test. It costs nothing except the discipline of collecting the baseline first, which is the same discipline behind getting the measurement and data foundations in place before the build.

The exposure sitting outside procurement

Two things compound this. The first is workplace AI use running through personal accounts on consumer plans, which puts it outside IT's visibility and outside procurement's field of view entirely. Nobody can quote a credible share, and the argument doesn't need one: you cannot instrument a workflow you don't know exists. And the people using it that way are typically the capable and conscientious ones, because the hook is competence-adjacent: better first drafts, faster answers. Screening for vulnerability profiles will find the wrong population.

The second is drift, and here I'm arguing rather than reporting. Sustained exposure to one model's register plausibly pulls a professional's prose toward it even when no generated text is ever pasted anywhere, and there'd be no audit event to point at if it did. The stylistic version is mildly embarrassing: a house voice nobody chose. The expensive version is analytic. Teams stop generating the third option, the one the model wouldn't have suggested, because the first two arrived so fast that the search felt finished, and nothing in the output flags what was never considered.

All of which lands on the same place: the next budget round, where AI spend is defended with outputs per head and targets get reset against the resulting baseline. Renew on stage-one metrics and you commit next year's team to sustaining a volume that was only ever affordable because checking was skipped, and to paying down the rework at the same time. That's a measurement design problem before it's a tooling problem, and it's fixable this quarter rather than after the defects arrive.

Measure the step you made pleasant and you'll always be impressed. Measure the step you left tedious and you'll find out what the gain cost.

Questions people ask

What metrics actually show whether AI is degrading work quality?

Metrics attached to the verification half of the workflow: rejection rate at review, revision passes per artefact, proportion of load-bearing claims independently checked, discard rate, defect escape rate, and time to first externally reported defect. Volume metrics such as documents drafted or tickets closed cannot detect quality loss, because they measure the step AI accelerated. Read them in pairs: a volume metric rising while its paired review metric stays flat is the signal.

How long before AI-related quality problems show up in the numbers?

Later than the decisions they should inform, which is the structural trap. Generation metrics tend to move almost as soon as a tool lands, while defects surface through downstream consumers, customers and remediation work on a much longer lag. DORA's 2024 finding that delivery stability degraded as AI adoption rose is the closest thing to a general signal, and even that is an association across a survey population rather than a clock you can set by. Plan around the asymmetry: assume renewal and target-setting will happen inside the gap, and instrument leading indicators such as rejection and verification rates rather than waiting for lagging ones.

Should we ban personal AI accounts at work instead?

Expect bans to relocate the usage rather than remove it, and to destroy the visibility you need in the process. That's a prediction from how policy without an alternative usually goes, not a measured result. The more useful sequence is to provide a sanctioned route good enough that the personal one isn't worth the friction, then instrument the review step inside it. The exposure you can measure is manageable; the exposure running through consumer accounts is neither measured nor governed.

Related

Written by an AI editorial persona of Abyshire's proprietary editorial system and reviewed by our team.