Agent-assisted onboarding to a legacy codebase

This is a methodology scenario, not a client engagement. Chokmah has not delivered this engagement for a named client. The workflow, constraints and sequence describe how we would run it. Any figures shown are illustrative targets, not measured results.
Agent-assisted onboarding to a legacy codebase is often assumed to speed new joiners up, but the one rigorous RCT found experienced developers 19% slower in mature codebases while believing they were faster. This page describes how Chokmah would measure the effect for a specific team before recommending any rollout.
- METR's RCT found experienced devs 19% slower in mature codebases, believing they were 20% faster.
- METR has since caveated the finding as historical; both facts belong on the table.
- Perceived speed-up is not evidence; measure task outcomes before rolling out.
- Lines and commits are vanity metrics an assistant trivially inflates.
- The result is a dated measurement of one team, not a verdict on AI coding tools.
The situation
An engineering team in a product GCC maintains a large, mature codebase that new joiners take months to become productive in. The reflex is to roll out an AI coding assistant and expect a speed-up. This scenario is about testing that assumption before acting on it, not assuming it.
The one rigorous randomised controlled trial available on this question found the opposite of the expectation. METR's 2025 study of experienced open-source developers working in large, mature codebases found they were 19% slower with early-2025 AI tools, while believing they had been sped up by about 20% (METR, July 2025). METR has since noted, in a February 2026 update, that this finding is historical and does not necessarily reflect current tools or workflows (METR, February 2026). Both facts belong on the table: the rigorous result was negative for exactly this kind of codebase, and the tools have moved since.
The honest position is to measure this team, in this codebase, before rolling anything out. This scenario is a composite of agent-assisted engineering onboarding as it appears in product teams. It names no client and reports no achieved result.
Constraints
- The only rigorous RCT on this question found a negative effect on experienced developers in mature codebases. That must be confronted, not ignored.
- Self-reported productivity is unreliable here. Developers in the METR study believed they were faster while being slower: perceived speed cannot be the measure.
- A large mature codebase carries implicit context an assistant does not have. Confidently wrong suggestions cost more than no suggestion in this setting.
- Onboarding productivity is genuinely hard to measure. Naive proxies like lines or commits reward the wrong behaviour.
- The tooling landscape changes fast. Any finding is dated the moment it is produced and must be framed as a measurement, not a verdict.
How we would run it
- 1Week 1
Define what onboarding productivity actually means here, with the team
Before any tool question, agree what a productive new joiner does in this codebase and how you would know. That usually means time to a first safely merged non-trivial change, ability to locate the right code for a task, and rework rate on early changes, not lines written or commits made. The definition is co-authored with the engineering lead, because a measure the team does not accept produces a result nobody trusts.
- 2Week 1
Confront the METR finding openly with the sponsor
Put the negative RCT and its caveat in front of the sponsor at the start. The point is not to argue for or against assistants; it is to establish that perceived speed-up is not evidence, that the rigorous result was negative for exactly this codebase type, and that the tools have since changed. Setting this up front is what makes an honest measured rollout possible instead of a faith-based one.
- 3Week 2
Design a measured comparison, not a feelings survey
Structure a comparison that avoids the trap METR identified: measure task outcomes and time on realistic onboarding tasks rather than asking developers how they felt. Where a controlled comparison is feasible, use one; where it is not, use before-and-after on matched task types with the same person. Pre-register the metric and the analysis so the result cannot be reinterpreted after the fact to suit whichever answer is wanted.
- 4Weeks 3-4
Run the assistant against real onboarding tasks and record outcomes
Have new and near-new joiners perform realistic tasks in the codebase with and without assistance, recording time, whether the change was correct, and how much rework it needed. Capture where the assistant helped (boilerplate, unfamiliar APIs, test scaffolding), and where it hurt: confidently wrong navigation in code with implicit context. The mixed picture is the finding, and it is more useful than a single verdict.
- 5Weeks 4-5
Analyse against the pre-registered metric and separate perception from outcome
Report the measured outcome against the metric agreed in week one, and separately report what participants believed, so the gap between the two is visible. If the tools help this team in this codebase, the data will show it and the rollout is evidence-based. If they do not, the data prevents an expensive mistake. Either way the honest number is the deliverable.
- 6Week 6
Deliver a rollout recommendation with the boundary conditions stated
The recommendation is specific: which task types and which parts of the codebase the assistant helps with, which it does not, and what to measure continuously if it is adopted. It is dated and caveated, because the tooling will change and this is a measurement of a moment, not a permanent law. The team keeps the task set and the measurement method so the question can be re-run when the tools move.
What we would not automate
Rolling out an assistant on the strength of perceived speed-up
The METR study's central lesson is that experienced developers felt faster while being measurably slower. Rolling out on self-report repeats exactly the error the one rigorous trial exposed. We would refuse to recommend adoption on perception, and would insist on a measured outcome against a pre-agreed metric before any team-wide decision.
Measuring onboarding with lines of code, commits or acceptance-rate vanity metrics
These proxies reward volume and suggestion-acceptance rather than correct, maintainable change, and they are trivially inflated by an assistant. Optimising them produces more code and worse onboarding. The measure has to be task outcomes and rework, even though those are harder to collect, because the easy metrics point the wrong way.
Presenting the result as a permanent verdict on AI coding tools
The tools change fast and METR itself has caveated its finding as historical. A responsible output is a dated measurement of this team in this codebase with this generation of tooling, plus a method to re-run it, not a claim that AI assistants help or hurt in general. Overclaiming in either direction would be dishonest.
Illustrative targets
These are targets used to frame the engagement, not measured results from a delivered client project.
Illustrative target: The onboarding-productivity measure itself: Pre-registered, agreed up front
The deliverable is a trustworthy measurement, so the metric (time to a first safely merged non-trivial change, and rework rate on early changes) is defined and pre-registered with the team in week one, before any tool is run, to prevent reinterpretation afterwards.
Illustrative target: Perception-versus-outcome gap: Measured and reported explicitly
Because the METR result turned on developers believing they were faster while being slower, both the measured outcome and the participants' self-reported experience are recorded, and the gap between them is reported rather than hidden.
Illustrative target: Task types where assistance helps or hurts: A boundary map, not one number
The honest result is mixed: help on boilerplate and unfamiliar APIs, harm on navigation in code with implicit context. The output is a map of where the assistant helps and where it does not for this codebase, which is more decision-useful than one aggregate figure.
Frequently asked questions
The one rigorous randomised trial says not reliably, for this exact case. METR's 2025 study found experienced developers were 19% slower with early-2025 tools in large mature codebases, while believing they were about 20% faster. METR has since noted the finding is historical and may not reflect current tools. The honest answer is to measure your team in your codebase before rolling out, rather than assume a speed-up that self-report will falsely confirm.
Because self-report is exactly what failed in the METR study: developers felt faster while being measurably slower. Perceived productivity and measured productivity diverged. So the engagement measures task outcomes and rework on realistic onboarding tasks against a metric agreed before anything is run, and reports the gap between what was measured and what participants believed.
Recommend a rollout on the strength of perceived speed-up; measure onboarding with lines of code or commits, which an assistant inflates while onboarding gets worse; or present the result as a permanent verdict on AI coding tools. The tools change fast, so the output is a dated measurement of one team with a method to re-run it, not a general claim.
This is Adoption Diagnostic work: a measurement engagement rather than a build. It runs over roughly the same two-to-six-week window, defining the metric with the team, confronting the evidence, running a measured comparison, and delivering a rollout recommendation with its boundary conditions and its date stated. The team keeps the task set and the method.
No. It is a methodology scenario, not a case study. Chokmah has not delivered this for a named client. The workflow is a composite of how agent-assisted engineering onboarding appears in product teams, and every figure is an illustrative target or a metric definition, not an achieved result.
Have a workflow like this?
We name the workflow before we start. Book a free AI Reality Check and build one live.