Point of view
AI made senior developers 19% slower, and they believed it sped them up
The one rigorous randomised trial of AI coding assistants found experienced developers got slower, not faster, while feeling faster. You cannot manage a productivity claim you have not measured, because the people doing the work are not reliable witnesses to their own speed.

METR's randomised controlled trial (16 experienced developers, 246 real tasks on mature codebases, early-2025 tools) found developers took 19% longer with AI while estimating they were about 20% faster. METR now says the result is historical and does not necessarily reflect current tools.
- METR, July 2025: experienced developers were 19% slower with early-2025 AI tools in a randomised trial.
- The same developers believed AI sped them up ~20%: a ~39-point gap between belief and measurement.
- METR stated in February 2026 that the result is historical and may not reflect current tools; the interval was +2% to +39%.
- Mature, context-heavy codebases are the hard case: exactly the enterprise condition.
- Self-reported productivity is not data; baseline before you roll out and pre-commit to kill criteria.
The evidence
In a randomised controlled trial, experienced open-source developers took 19% longer to complete real tasks when allowed to use early-2025 AI tools.
The same developers estimated afterwards that AI had made them about 20% faster: a gap of roughly 39 percentage points between belief and measurement.
METR now states the early-2025 result is historical and does not necessarily reflect current tools or workflows; the confidence interval ran from +2% to +39%.
The trial covered 16 experienced developers working across 246 real tasks on large, mature repositories they already knew well.
What did METR actually test?
In a randomised controlled trial, 16 experienced open-source developers worked through 246 real tasks on large repositories they already knew well; those permitted to use early-2025 AI tools took 19% longer than those who were not (METR, July 2025). This is, as far as we know, the only randomised trial of frontier coding assistants on mature codebases published to date. Almost everything else you have read about AI and developer productivity is a self-report survey or a vendor benchmark.
The design matters because it removes the usual confound. These were not junior developers on a greenfield toy project: the setting where assistants shine. They were experts on code they understood, which is the setting most enterprise engineering actually is. And the direction was negative.
We put this number in front of every funded SaaS and product team we talk to, because it is the audience most likely to have rolled out Copilot to the whole engineering org on the strength of a demo, with no baseline, and no way to know whether it helped.
Experienced developers took 19% longer with AI tools, while estimating afterwards that AI had sped them up by about 20%. >METR, early-2025 developer productivity study (2025)
METR stated in February 2026 that this result is historical and does not necessarily reflect current tools or workflows, with a confidence interval from +2% to +39%. We cite the caveat every time we cite the number. A statistic quoted without its author's own correction is exactly the selective evidence we sell against.
Why is the 39-point gap the real finding?
Strip out the specific 19% and one result survives every caveat: the developers were wrong about their own speed, and wrong by a wide margin. They measured slower and felt faster: a gap of roughly 39 percentage points between belief and measurement (METR, arXiv:2507.09089, 2025).
This is the part that generalises even if the tools have since improved. AI assistance feels productive. It removes the friction of the blank page, produces fluent output immediately, and keeps you in a state of continuous forward motion. None of that is the same as finishing sooner. The felt experience of being unblocked is not evidence of throughput, and human beings are structurally bad at telling the two apart.
Which means self-reported productivity is not data. If your entire case for an AI rollout is that the team says it helps, you have the exact instrument the trial showed to be unreliable. This is why we baseline before we build, and why we distrust every dashboard that measures how the tool feels rather than what the work now costs: the distinction we call workflow absorption.
Why are mature codebases the hard case?
The assistant's weakness is context. On a fresh project, the entire relevant world fits in the model's window and the suggestions are good. On a ten-year-old repository with local conventions, implicit invariants, and a reason step four exists that nobody wrote down, the assistant confidently proposes code that is plausible and wrong. The expert then spends time reading, verifying and correcting: often longer than writing it themselves would have taken.
That is the enterprise condition. A GCC's most valuable engineering is precisely the deep, context-heavy work on systems the organisation has run for years. It is the least AI-tractable and the most tempting to automate, because it is where the senior salaries sit.
The lesson is not that assistants are useless. It is that where they help and where they hurt is an empirical question about your codebase, not a property of the tool. An engagement scenario for onboarding a legacy codebase is exactly the shape of work where we would measure first and generalise never.
What does METR itself say about the result now?
In February 2026, METR published an update stating that the early-2025 finding is historical and does not necessarily reflect current tools or workflows, and it is redesigning the experiment to control for selection effects it observed in follow-up work (METR, 24 February 2026). The original confidence interval ran from +2% to +39%: wide, and entirely on the slower side of zero.
We want to be scrupulous here, because this caveat is the whole point. Do not use this page to argue that AI makes developers slower in 2026. That is not what the current evidence says, and METR has explicitly disowned that reading of the number. The tools of mid-2026 are not the tools of early 2025.
What the evidence still supports is narrower and more durable: a rigorous measurement of experienced developers found a negative effect where everyone expected a positive one, and the developers could not feel it. That should make you humble about any productivity claim (up or down) that has not been measured on your own team, on your own code.
What does this mean for a transformation owner?
Measure before you roll out. Not a survey: a measurement. Pick a representative slice of work, record what it costs now in cycle time and rework, then run a controlled comparison rather than asking people how it went. If a rollout is worth the licence spend, it is worth two weeks of baselining first.
Write the kill criteria in advance. Decide, before the pilot, what result would make you stop. A team that has pre-committed to a number is immune to the perception gap; a team that will judge the pilot on how it felt is not.
And resist the org-wide default rollout. The METR result is an argument for targeting, not for abstinence. Different work, different codebases and different seniority levels will land in different places, and the only way to know is to instrument it. That is what a two-week Adoption Diagnostic produces before anyone commits budget, and it is why we build measurement into every evaluation harness we ship.
Disclosure: Chokmah is a new practice with no completed client engagements. The argument above rests on published research and method, not on results we are claiming.
Measure your team before you scale the licence
A two-week diagnostic baselines the work you are about to hand to an assistant, so the productivity claim is a number you can defend to a CFO, not a feeling.
Book an Adoption Diagnostic · Read how we measure the workflow
Frequently asked questions
Did METR find AI makes all developers slower?
No. It found a 19% slowdown for 16 experienced developers working on large, mature codebases they already knew well, using early-2025 tools. It did not test junior developers, greenfield projects or 2026 tools, and METR now says the result is historical and should not be read as a claim about current tools.
How many developers were in the METR study?
Sixteen experienced open-source developers, working across 246 real tasks on repositories they maintained. It was a randomised controlled trial (developers were randomly assigned whether they could use AI on each task), which is what makes it stronger evidence than a self-report survey.
Is the METR result still valid in 2026?
METR stated in February 2026 that the early-2025 finding is historical and does not necessarily reflect current tools or workflows. Treat the specific 19% as a snapshot of early-2025 tools. The durable finding is the perception gap: measured and felt productivity pointed in opposite directions.
Why did developers think they were faster?
AI assistance removes the friction of starting and keeps you in continuous forward motion, which feels productive. Feeling unblocked is not the same as finishing sooner, and people are poor witnesses to their own speed. That is why self-reported productivity is not a reliable measurement instrument.
Does this apply to junior developers?
The study did not test them, so we do not know. There is a plausible case that assistants help more where the developer has less context to lose: greenfield work, unfamiliar languages, boilerplate. That is a hypothesis to measure on your own team, not a result to assume.
Key terms
- Workflow absorptionWorkflow absorption measures whether AI has actually changed how work runs (steps redesigned, cycle time reduced, errors cut) as opposed to adoption, which only counts access such as seats and logins.
- Evaluation harnessAn agent evaluation harness is a repeatable test suite that scores an AI agent's outputs against fixed, versioned cases before and after every change, so teams can tell regression from variance.
More points of view
- Measure the workflow, not the loginsLicence logins, completions and prompt counts measure that a tool was opened, not that the work changed. The 95% failure rate is a measurement failure as much as an execution one: organisations counted adoption and never instrumented absorption.
- Why 95% of GenAI pilots fail, and what the surviving 5% did differentlyThe 95% figure is contested and imperfect, and it still describes your pilot. The failure is not model quality. It is that nobody instrumented the workflow the tool was supposed to change, so no result could ever have been measured.
Frequently asked questions
No. It found a 19% slowdown for 16 experienced developers working on large, mature codebases they already knew well, using early-2025 tools. It did not test junior developers, greenfield projects or 2026 tools, and METR now says the result is historical and should not be read as a claim about current tools.
Sixteen experienced open-source developers, working across 246 real tasks on repositories they maintained. It was a randomised controlled trial (developers were randomly assigned whether they could use AI on each task), which is what makes it stronger evidence than a self-report survey.
METR stated in February 2026 that the early-2025 finding is historical and does not necessarily reflect current tools or workflows. Treat the specific 19% as a snapshot of early-2025 tools. The durable finding is the perception gap: measured and felt productivity pointed in opposite directions.
AI assistance removes the friction of starting and keeps you in continuous forward motion, which feels productive. Feeling unblocked is not the same as finishing sooner, and people are poor witnesses to their own speed. That is why self-reported productivity is not a reliable measurement instrument.
The study did not test them, so we do not know. There is a plausible case that assistants help more where the developer has less context to lose: greenfield work, unfamiliar languages, boilerplate. That is a hypothesis to measure on your own team, not a result to assume.
Bring the evidence to your team
We walk in with the failure rates, then the method. Book a free AI Reality Check.