Does AI Actually Make You More Productive?
Sometimes, and usually by less than people expect. Randomised experiments on short, self-contained tasks find gains of roughly 25 to 40%. Studies following people through their real jobs find effects that are real but narrower, concentrated in specific activities rather than spread across the week, and self-reported surveys of whole working weeks put the saving at a few percent of hours. One 2025 trial found experienced developers 19% slower while believing they had been faster. All of these are credible results measuring different things.
Where the gains are largest
Shakked Noy and Whitney Zhang, economists at MIT, gave 453 college-educated professionals two writing assignments of 20 to 30 minutes resembling their real work: press releases, short reports, awkward emails. Half were given ChatGPT, which at the time ran on GPT-3.5, in sessions held between 27 January and 21 February 2023. That group finished about 40% faster and scored 18% higher on quality, graded blind by experienced people in the same occupations, and the gap between strong and weak writers narrowed. Science published the study in July 2023.
Fabrizio Dell'Acqua and colleagues tested 758 Boston Consulting Group consultants in a field experiment published in Organization Science in March 2026. Across 18 realistic business tasks inside GPT-4's frontier, ranging from creative to analytical, consultants with AI completed 12.2% more of them and worked 25.1% faster, with human graders scoring their work significantly higher and the weakest performers gaining most. The researchers then set a task designed to sit just outside the model's competence. There, consultants using AI were 19 percentage points less likely to be right than the control group. The authors called that boundary the jagged technological frontier: capability is uneven in ways you cannot see from outside.
What happens in real jobs
Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied 5,172 customer support agents given an AI assistant, published in the Quarterly Journal of Economics in April 2025. Issues resolved per hour rose 15% on average, but the average hides the finding that matters. Less experienced and lower-skilled agents improved both the speed and the quality of their output, the novices by around 30%, while the most experienced and highest-skilled saw small gains in speed and small declines in quality. The tool appeared to transmit what good agents already knew to agents who did not, and to get slightly in the way of the people who already knew it. Three field experiments run by Zheyuan Cui and colleagues, covering 4,867 developers at Microsoft, Accenture and a Fortune 100 manufacturer, found completed weekly tasks up 26.08%, with a standard error of 10.3 percentage points, so the true figure could be a good deal smaller or larger than the headline. Junior developers again gained most.
A third study is useful because its results are modest. Eleanor Dillon, Sonia Jaffe, Nicole Immorlica and Christopher Stanton tracked 7,137 knowledge workers at 66 firms for six months after Microsoft 365 Copilot was randomly assigned, using telemetry rather than self-reports. Email time fell about 1.4 hours a week, while meeting time, documents completed and collaborative output did not move. What a person controls alone changed; what needs colleagues to change did not. Each of these studies measured how the work got done rather than how many people were doing it, which is a separate question our article on the jobs AI will replace takes on directly.
The trial that found the opposite
In July 2025 METR, a nonprofit research organization, published a randomised controlled trial by Joel Becker, Nate Rush, Beth Barnes and David Rein. Sixteen experienced open-source developers, with roughly five years on the repositories concerned, worked through 246 real issues in mature codebases of over a million lines. Each issue was randomly assigned to allow or forbid AI, mostly Cursor Pro with Claude 3.5 and 3.7 Sonnet, between February and June 2025.
Developers took 19% longer where AI was allowed, on a confidence interval running from 2 to 39% longer. Beforehand they had expected it to make them 24% faster; afterwards, having been measurably slower, they still estimated it had made them 20% faster. Economics experts forecast a 39% speed-up, machine learning experts 38%. The authors were careful about scope: sixteen developers is a small sample, the work met the review standards of projects with tens of thousands of GitHub stars, and participants had limited prior hours with the tools, so learning effects may not have been exhausted.
In February 2026 METR said it was redesigning the follow-up experiment. Developers were refusing to take part because they did not want to work without AI, which strips the heaviest adopters from the sample, and 30 to 50% of participants avoided submitting tasks they thought AI would handle well. METR puts that selection problem down primarily to its developers now expecting more uplift from AI, with a cut in pay from 150 dollars an hour to 50 also likely contributing. Its raw estimates from the later round run the other way from the original: 18% faster for the ten developers who returned, on a confidence interval from 38% faster to 9% slower, and 4% faster for the newly recruited ones, from 15% faster to 9% slower. METR's own reading is that developers are probably more sped up in early 2026 than they were in early 2025, that the missing developers and missing tasks are the high-uplift ones so these figures are likely a lower bound, and that the data is only very weak evidence for the size of that increase. The original result was not retracted. The method was judged unable to carry more weight, which is a different and more useful statement.
What each study measured
| Study | Setting and sample | Measured | Result |
|---|---|---|---|
| Noy and Zhang, 2023 | Lab tasks, 453 professionals | Time, blind-graded quality | 40% faster |
| Dell'Acqua et al. | Firm experiment, 758 consultants | Speed, graded quality | 25% faster, 19 points worse off-frontier |
| Brynjolfsson et al. | Live jobs, 5,172 support agents | Issues per hour | 15%, about 30% for novices |
| Cui et al. | Live jobs, 4,867 developers | Tasks per week | 26%, standard error 10.3 points |
| Dillon et al. | Live jobs, 7,137 office workers | Time in email, meetings, docs | Email down, collaboration flat |
| METR, 2025 | Open-source repos, 16 developers | Time per issue | 19% slower |
Feeling faster and being faster
The gap between perceived and measured speed is the most consistent finding here, and it appears at every level. METR's developers misjudged their own performance by nearly 40 percentage points while being timed. A Federal Reserve Bank of Atlanta paper from March 2026, surveying 748 corporate executives between November 2025 and January 2026, found reported AI productivity gains for 2025 of 1.8% against 0.6% implied by those firms' own revenue and employment figures.
Part of the explanation is that AI moves effort rather than removing it. BetterUp Labs and Stanford Social Media Lab surveyed 1,150 US desk workers in September 2025 and found 40% had received what they called workslop that month: output that looks finished but pushes the thinking onto whoever receives it, taking about two hours each to resolve. The sender saved time. The organization did not.
Writing with AI is fast while reading it is not, so the moment that feels productive is not where the time goes. Verification is the expensive part and easy to skip, which is automation bias in action. It matters most where the tool can be confidently wrong, as our article on AI hallucinations explains.
Why economy-wide numbers lag
None of this is new. In July 1987 the economist Robert Solow wrote in the New York Times Book Review that "you can see the computer age everywhere but in the productivity statistics." Computers were visibly changing offices while national figures barely moved, and the gains arrived later, once firms reorganised around the technology rather than bolting it on. A brief published by the International Labour Organization in May 2026 gives the current version of the AI productivity paradox a sharper name, the aggregation paradox: task-level gains of 10 to 70% that shrink at firm level and disappear at sector and economy level. Anders Humlum and Emilie Vestergaard, who linked adoption surveys of 25,000 Danish workers across 7,000 workplaces to national administrative records, find effects on earnings and recorded hours precise enough to rule out anything larger than 2% two years after ChatGPT launched, even while the structure of the work underneath those figures shifts. Reported time savings are small enough to be consistent with that: about 2.8% of work hours among Danish chatbot users, and roughly 1.4% in a US survey in late 2024.
Why the studies disagree
Drafting a press release and fixing a subtle bug in code you have maintained for five years are not the same activity, and the jagged frontier means the difference is not always visible from the outside, which our article on what AI cannot do sets out in more detail. Every field study that examined experience found bigger gains for less experienced workers, so the average depends on who is in the sample. Speed, quality, output count and earnings can move in opposite directions at once. Models change faster than research publishes, and each study carries its own selection problems.
A fair summary: gains are well established for short, text-heavy, checkable tasks, especially for people newer to the work; real but smaller in live jobs; unproven or sometimes negative for experienced specialists working on complex material to high correctness standards.
What this suggests for your own work
The evidence is strongest where you are producing a first version of something textual you would review anyway, where you can judge the result, and where being wrong is recoverable. It is weakest where you are the expert, the material is complex, correctness matters, and checking costs as much as producing. Our guide on how to start using AI well covers the habits that help, and writing better prompts moves more of the work into the part the evidence actually supports.
Three things are worth watching. Measure time from starting to something being accepted, not time to a first draft, because the draft is where the sense of speed lives. Track rework, both yours and what you create for colleagues. And notice what happens to skills you stop practicing, which is deskilling and shows up months later.
The honest position is that nobody yet knows the size of the effect for your particular work, and the people best placed to find out are the ones doing it. Treat the published figures as a guide to where to look rather than a number to expect, and keep a record of your own. It will be more reliable than your memory of last week, which is the one measurement this whole body of evidence consistently finds wrong.
Related AI terms
- AI productivity paradox: The gap between the gains people expect from AI and the gains they see.
- Workslop: Poor AI-created work that leaves colleagues to work out what it means or fix it.
- Automation bias: Trusting a computer's suggestion too much, even when other evidence says otherwise.
- AI augmentation: Using AI to extend a person's capabilities while they stay involved in the work.
- Automation paradox: Automation cutting everyday human involvement while making rare problems harder.
- Deskilling: The loss of skills when people no longer practice them.
Frequently Asked Questions
Has AI actually increased productivity?
At the level of an individual task, yes, and the effect is well measured. At the level of a national economy, it has not shown up yet. Randomised experiments on short self-contained tasks find gains of roughly 25 to 40%, and field studies in real jobs find real but narrower effects. National statistics two years after ChatGPT launched are precise enough to rule out anything larger than 2% in Danish earnings and recorded hours. Both can be true at once. Task-level gains are diluted at team level, diluted again at firm level, and vanish into the noise of an economy, which is the pattern Robert Solow described about computers in 1987 and which economists now call the aggregation paradox.
How does AI make you more productive?
Where it works, it works by removing the slowest part of a task rather than by making you generally faster. The gains concentrate in first drafts, boilerplate, summarizing and restructuring: work where producing something is the bottleneck and checking it is cheap. They shrink or reverse where checking costs as much as producing, which is why the largest measured gains went to writing and customer support and the one measured slowdown went to experienced developers working in code they already knew well. The other consistent finding is who benefits: in every field study that looked, less experienced people gained the most and the most experienced gained least.
How should a team actually measure whether AI is helping?
Copy the design of the better studies. Compare against a group doing similar work without the tool, because without a comparison you are measuring the quarter rather than the tool. Use system records rather than recall, which is why the Microsoft 365 study of 7,137 workers used telemetry. Measure at the level of finished work that someone accepted, not a first draft. Run it long enough to see rework appear, since rework usually lands on a person other than the one who felt fast. And report the spread as well as the average, because in every field study here the average hid a large difference between experienced and inexperienced people.
Do the gains fade once the novelty wears off?
The evidence available says they hold rather than fade, though none of it runs very long. In the customer support study the increase in issues resolved per hour appeared in the first month of deployment, grew slightly in the second, and then stayed stable for the rest of the sample. Agents who followed the AI's suggestions closely kept part of the improvement during outages when the tool was unavailable, which points to learning rather than dependence, while agents who ignored the suggestions kept nothing. What has not appeared, two years after ChatGPT launched, is those steady individual gains showing up in earnings or recorded hours.
My employer is quoting a vendor's productivity figure at me. How should I read it?
Ask four things. What exactly was measured, since time to a first draft and time to work a colleague accepted are very different numbers. Who was in the sample, because every field study here found larger gains for less experienced people, so a figure from a junior team will not transfer to a senior one. Was there a comparison group who did the same work without the tool. And when was it measured, with which model, because the tool and the moment determine the number. A figure that cannot answer those four questions is a marketing claim rather than a measurement.
Does this research apply to the AI tools available now?
Only loosely, and that is worth saying out loud. Every study here tested a named tool in a named window: GPT-3.5 in early 2023, GPT-4 later that year, Cursor with Claude 3.5 and 3.7 Sonnet in the first half of 2025. METR's own follow-up work on later tools suggests developers are probably more sped up now than they were in early 2025, but METR is clear that its data is only very weak evidence for how much. Treat any specific percentage as a reading from one tool at one moment rather than a fixed property of AI.
Should I stop using AI if I am the expert on the material?
The evidence does not support stopping, and it does not support assuming you are faster either. It supports being selective. The tasks where experienced specialists lose time are the ones where the material is complex, the correctness bar is high, and checking the output costs about as much as producing it yourself. Those are worth keeping. The work around them, the first drafts and the summaries and the things you would review anyway, is where the case is strongest and where being wrong is cheapest to fix.
Sources
- Shakked Noy and Whitney Zhang, "Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence," published in Science, 13 July 2023; author copy. https://shakkednoy.com/Noy%20Zhang%20NBER%20SI.pdf
- Fabrizio Dell'Acqua, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon and Karim R. Lakhani, "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality," Organization Science, Articles in Advance, pages 1 to 21, published online 11 March 2026, DOI 10.1287/orsc.2025.21838; published version hosted by Harvard Business School. https://www.hbs.edu/ris/Publication%20Files/dell-acqua-et-al-2026-navigating-the-jagged-technological-frontier_5c589c8c-fbb5-458f-b285-c944746cd717.pdf
- Erik Brynjolfsson, Danielle Li and Lindsey Raymond, "Generative AI at Work," Quarterly Journal of Economics 140(2), April 2025, pages 889 to 942; published abstract and citation record, the publisher page itself returning 403. https://ideas.repec.org/a/oup/qjecon/v140y2025i2p889-942..html
- Zheyuan Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng and Tobias Salz, "The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers," February 2025. https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf
- Eleanor Dillon, Sonia Jaffe, Nicole Immorlica and Christopher Stanton, "Shifting Work Patterns with Generative AI," NBER Working Paper 33795, May 2025, revised November 2025. https://www.nber.org/papers/w33795
- Joel Becker, Nate Rush, Beth Barnes and David Rein, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," METR, 10 July 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- Joel Becker, Nate Rush, Beth Barnes and David Rein, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," arXiv:2507.09089, July 2025. https://arxiv.org/abs/2507.09089
- METR, "We are Changing our Developer Productivity Experiment Design," 24 February 2026; source of the +2 to +39 percent confidence interval on the original result, the selection analysis, and the later-round estimates. https://metr.org/blog/2026-02-24-uplift-update/
- Salome Baslandze, Zachary Edwards, John R. Graham, Ty McClure, Brent Meyer, Michael Sparks, Sonya Ravindranath Waddell and Daniel Weitz, "Artificial Intelligence, Productivity, and the Workforce: Evidence from Corporate Executives," Federal Reserve Bank of Atlanta Working Paper 2026-4, March 2026. https://www.atlantafed.org/-/media/Project/Atlanta/FRBA/Documents/research/publication/working-paper/2026/03/25/04-artificial-intelligence-productivity-and-the-workforce-evidence-from-corporate-executives.pdf
- BetterUp Labs and Stanford Social Media Lab, "Workslop: The Hidden Cost of AI-Generated Busywork," survey of 1,150 US desk workers, September 2025. https://www.betterup.com/workslop
- Kate Niederhoffer, Gabriella Rosen Kellerman, Angela Lee, Alex Liebscher, Kristina Rapuano and Jeffrey T. Hancock, "AI-Generated Workslop Is Destroying Productivity," Harvard Business Review, 22 September 2025. https://hbr.org/2025/09/ai-generated-workslop-is-destroying-productivity
- Robert Solow, "We'd Better Watch Out," New York Times Book Review, 12 July 1987, quoted by the Reed College Department of Economics. https://www.reed.edu/economics/parker/201/cases/it_productivity.html
- Cheuk Yu Cheryl Chan and Khatia Shedania, "The Aggregation Paradox of AI: Why do micro-economic productivity gains from AI disappear at scale," International Labour Organization, 6 May 2026. https://www.ilo.org/publications/aggregation-paradox-ai-why-do-micro-economic-productivity-gains-ai
- Anders Humlum and Emilie Vestergaard, "Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI," NBER Working Paper 33777, May 2025, revised March 2026, previously circulated as "Large Language Models, Small Labor Market Effects." https://www.nber.org/papers/w33777
- Alexander Bick, Adam Blandin and David Deming, "The Rapid Adoption of Generative AI," NBER Working Paper 32966, September 2024, revised February 2025. https://www.nber.org/papers/w32966
- Erik Brynjolfsson, Danielle Li and Lindsey R. Raymond, "Generative AI at Work," NBER Working Paper 31161, April 2023, revised November 2023; working paper full text, used for the month-by-month pattern of effects and the AI outage analysis. https://www.nber.org/system/files/working_papers/w31161/w31161.pdf
- METR, "Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity," 11 May 2026. https://metr.org/blog/2026-05-11-ai-usage-survey/