What Those Headlines Actually Measured
On one screen: “Generative AI increases productivity by 66% on average.” On the other: “95% of enterprise AI pilots fail to produce measurable ROI.” Both statements are cited as fact in the same year. If you are a knowledge worker trying to decide how to use AI for personal productivity, which number do you believe? I look at these two numbers and think: the useful question is not which one is right — it’s what each one actually measured.

The 66% comes from a NN Group study that measured task completion time across three occupations. The 95% failure rate is from the MIT NANDA report, which defines “failure” as not reaching measurable ROI — not as “AI does nothing.” Both are real. The task-level gains are genuine for specific jobs; the enterprise failure rate measures organizational adoption, not individual output. The gap between them is where the real story lives. A Forbes synthesis of dozens of studies puts the individual gain range at 14–55%. That is not one number. It is a range that depends on task type, worker skill, and measurement method.
The NN Group breakdown shows the variation clearly:
| Occupation / Task | Productivity Increase |
|---|---|
| Customer support agents | 14% |
| Business document writers | 59% |
| Programmers (coding) | 126% |
These are task completion times under controlled conditions. They measure output quantity, not quality — though the NN Group did include a quality check. The St. Louis Fed found a narrower effect: workers save 5.4% of their hours overall, but gain 33% per hour of actual AI use — because they are not using it every hour. The OECD reports a sector range of 5% to over 25%. What this means for personal productivity: AI helps enormously on tasks that are repetitive, well-defined, and have clear output checks. It helps far less on open-ended reasoning, nuanced communication, and tasks that require deep domain judgment. The aggregate average is misleading. The relevant number is the one that matches your actual work.
The Novice Gains, The Expert Slows Down

The most counterintuitive finding is that AI does not amplify expertise — it compresses it. Erik Brynjolfsson's study of 5,179 customer service agents found that lower-skilled workers improved their resolution rate by 34%, while top performers showed minimal gains and sometimes quality declines. But I need to be careful here: that study comes from a narrow, scripted domain — customer support. It does not mean the same pattern automatically applies to complex knowledge work like strategy or product design. Still, the logic is plausible: novices benefit because AI fills gaps in their knowledge; experts have well-established workflows, and the AI's suggestions often conflict with their mental model. They spend time verifying, correcting, and deciding when to override. That takes real time.
The Perception-Reality Gap
A METR study of 16 experienced developers found that participants took 19% longer to complete real coding tasks with AI, yet believed they were 20% faster — a 39-point gap between perception and reality. I would not treat 16 developers as a population estimate. It is a case study, and a small one. But the pattern replicates across other studies. When you are highly skilled in a domain, AI does not make you faster. It makes you doubt, verify, and adjust. That takes time.
The clearest demonstration of AI's capability boundary comes from the BCG/Harvard jagged frontier study of 758 consultants. Within AI's capability boundary, users completed 12% more tasks, 25% faster, with 40% higher quality. But on tasks just outside that boundary, they were 19 percentage points more likely to produce incorrect solutions. The frontier is jagged — not a smooth line. The consultants could not reliably predict which tasks were safe. That is the trap: misplaced trust. The Klarna case — handling 2.3 million conversations in a month, resolution time dropping from 11 minutes to under 2 minutes — is a good example of a task well inside the frontier. Customer support scripts are structured and high-volume. That is not the same as strategic analysis.

The Costs That Do Not Show Up in Lab Studies
If AI's gains are real but uneven, why do so many pilots fail? The 95% figure from MIT NANDA measures “failure” as not delivering measurable ROI — not that AI does nothing. The hidden costs explain why the ROI does not appear:
- Rework: 41% of workers have encountered AI-generated “workslop” that required nearly two hours of correction per instance (Stanford research).
- Verification time: The experts in the METR study spent extra time checking AI output, erasing the time saved on generation.
- Perception gap: 92% of daily AI users report satisfaction in surveys, but objective measures tell a different story. If users think they are faster when they are not, they will not redesign their workflow to fix the problem.
These costs do not appear in controlled lab studies that measure pure task completion time. In real work, they compound. The 14–55% gains shrink if every output needs a review pass. That is the difference between a study and your actual Tuesday.
The Real Work Is Redesigning Work
BCG's 10-20-70 rule is a consulting heuristic, not a research finding. But it matches the evidence: 10% of value comes from the algorithm, 20% from technology and data, and 70% from people and process redesign. The Census Bureau data shows that only 5% of U.S. firms have meaningfully adopted AI. The gap between capability and practice is enormous.
For personal productivity, the 70% means: do not just add AI to your existing workflow. Change the workflow. Build a verification step after AI drafting. Create a checklist for tasks near the jagged boundary. Train yourself to recognize when AI output feels plausible but is likely wrong. That is where the real gain lives — not in the model, but in how you use it.
When to Use AI: A Simple Rule of Thumb
The evidence points to two key dimensions: your skill level relative to the task, and the task's location relative to AI's jagged frontier. A simple 2x2 matrix can guide your choice.
| Low skill | High skill | |
|---|---|---|
| Inside frontier | Delegate fully to AI | Use AI as a junior assistant; light review |
| Outside frontier | Use AI with heavy verification | Skip AI (or use it for inspiration only) |
The biggest gains come from pairing AI with human oversight and process redesign, not from adding more tools. If you do not redesign your workflow, the gains will be eroded by hidden costs. That is the fine print the headlines never mention.
Comments
Join the discussion with an anonymous comment.