Adoption
Adoption metrics are lying to you: the six numbers that are not
Logins, completions and prompt counts measure activity, not outcome. Six metrics that hold up under scrutiny, how to baseline a workflow in a week, and how to instrument without surveilling.


Licence logins, completions, and prompt counts measure activity, not outcome. The six metrics that hold up are cycle time per instance, rework rate, escalation rate, cost per successful task, share of steps redesigned, and time-to-competence. MIT found 95% of pilots produced no measurable P&L impact.
- Logins, completions and prompt counts are activity metrics; they do not predict outcomes.
- The six that survive scrutiny are cycle time, rework rate, escalation rate, cost per successful task, steps redesigned, and time-to-competence.
- MIT NANDA, 2025: 95% of pilots produced no measurable P&L impact: often because nobody baselined.
- Baseline the workflow in a week from data that already exists.
- Instrument the workflow, not the person: surveillance corrupts the inputs.
Licence logins, completion certificates, and prompt counts measure activity, not outcome. The six metrics that hold up are cycle time per workflow instance, error/rework rate, escalation rate, cost per successful task, share of workflow steps redesigned, and time-to-competence. MIT found 95% of pilots produce no measurable P&L impact largely because nobody baselined the workflow.
Key takeaways
- Logins, completions and prompt counts are activity; they do not predict outcomes.
- Six metrics survive scrutiny: listed below, all measured against a pre-baseline.
- MIT NANDA, 2025: 95% of pilots produced no measurable P&L impact.
- Instrument the workflow, not the person.
Why do adoption dashboards look good while nothing changes?
Because they measure the wrong verb. An adoption dashboard counts access (seats provisioned, tools opened, courses completed, prompts sent), and every one of those numbers rises when you push it, none of them requires the underlying work to change, and all of them are easy to pull. So the dashboard is green while cycle times are flat, and everyone is surprised at renewal.
The distinction is between adoption and workflow absorption. Adoption counts that a tool was available. Absorption counts that the work is now different. MIT's NANDA study found 95% of enterprise GenAI pilots produced no measurable P&L impact (MIT NANDA, July 2025), and a large part of that is a measurement failure as much as an execution one: organisations counted adoption and never instrumented absorption. We make the full argument in measure the workflow, not the logins.
The six metrics that survive scrutiny
Each of these counts a change in the work, not an activity, and each is measured against a pre-baseline.
| Metric | What it counts | Why it is not vanity |
|---|---|---|
| Cycle time per instance | Time from start to accepted output | The number the sponsor already believes |
| Error / rework rate | Share returned or corrected downstream | Catches fast-but-wrong AI output |
| Escalation rate | Share leaving the standard path | Detects work pushed sideways, not done |
| Cost per successful task | Total cost ÷ successful outcomes | The honest denominator; exposes retries |
| Steps redesigned | Share of workflow steps actually changed | The direct measure of absorption |
| Time-to-competence | How long a person takes to run the new workflow well | Predicts whether the change sticks |
Six numbers, one workflow. Notice what is absent: no logins, no completions, no prompt counts, no seat utilisation. Those are not on the list because none of them moves when the work stays the same, which is exactly the condition the six are built to detect.
How to baseline a workflow in one week
A week is enough because you are not building new instrumentation. You are reading data that already exists. Ticketing systems hold timestamps. Document stores hold versions and rework. Approval tools hold escalations. The week's work is to define the workflow's boundaries, pull those numbers for recent instances, and write them down before you change anything.
The non-negotiable is that the baseline is captured before the tool arrives. A baseline reconstructed afterwards is worthless, because memory moves in a predictable, generous direction. METR's randomised trial is the cleanest proof: experienced developers were measured as 19% slower using early-2025 AI tools on real tasks in mature codebases, while believing afterwards they had been sped up by about 20% (METR, July 2025). METR has since said the result is historical and does not necessarily reflect current tools (METR, February 2026). The tools improved; the gap between felt and measured performance is a property of people, and it is the entire reason you baseline before you build rather than surveying after.
Instrumenting without surveilling your team
There is a real risk in all of this: instrument the person instead of the workflow and you get surveillance, resentment, and corrupted data. Every one of the six metrics is a workflow-level number. None requires watching an individual's prompts or keystrokes.
Measure the queue, not the person. Cycle time is a property of the workflow instance, not of who touched it. Rework rate is counted at the point work is returned, not attributed to a name. When you report, report by workflow, never by individual. This is not only an ethical line; it is a data-quality line. The moment people believe the numbers will be used against them individually, they manage the numbers, and your baseline becomes theatre. The same principle governs an evaluation harness: it scores the system's runs, not the operators.
Reporting the number that might embarrass you
The final discipline is the hardest. Report the metric even when it is bad, and report it in the same format every time. A metric that appears only when it flatters the project is a highlight reel, and sponsors learn to discount it.
A workflow that came back worse after the intervention is information, not failure. It tells you to change the approach or stop, and both are cheaper now than at the production gate. The organisations that stay in the 95% are the ones whose dashboards only ever showed green. The ones that escape it are the ones that reported a red number in month two and acted on it.
What this means for a GCC transformation owner
You will be tempted to report the activity metrics, because they exist, they are easy, and they rise on command. That temptation is how programmes end up green all the way to cancellation. Report the six instead. They are harder to move, which is the point: a number that is hard to move is a number worth reporting, because when it does move you have proof.
Start with one workflow, baseline it this week, and put its cycle time in front of your sponsor next to the date the delta will be measurable. That is a smaller, truer claim than any adoption dashboard, and it is the claim that survives the renewal conversation. Building that measurement discipline is the core of what an adoption diagnostic installs.
Sources
- MIT NANDA, The GenAI Divide: State of AI in Business 2025, July 2025. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- METR, We are Changing our Developer Productivity Experiment Design, 24 February 2026. https://metr.org/blog/2026-02-24-uplift-update/
Related reading: measure the workflow, not the logins · what workflow absorption means · book an adoption diagnostic
Read next
- Measure the workflow, not the loginsLicence logins, completions and prompt counts measure that a tool was opened, not that the work changed. The 95% failure rate is a measurement failure as much as an execution one: organisations counted adoption and never instrumented absorption.
- Workflow absorptionWorkflow absorption measures whether AI has actually changed how work runs (steps redesigned, cycle time reduced, errors cut) as opposed to adoption, which only counts access such as seats and logins.
- Evaluation harnessAn agent evaluation harness is a repeatable test suite that scores an AI agent's outputs against fixed, versioned cases before and after every change, so teams can tell regression from variance.
More from the blog
- Writing an AI business case that survives CFO reviewAI business cases fail because they cannot be falsified. Name one workflow, state a measured pre-baseline, and write the kill criteria in advance. A one-page template.
- Your Copilot licences are unused. Here is the 30-day diagnostic.Low seat utilisation is a workflow problem wearing a training problem's clothes. A 30-day audit that finds out which, without surveilling your team.
- How to evaluate an AI agent in productionModel benchmarks do not transfer to your workflow. Score every run against versioned cases on four axes, build a golden set from real tickets, and catch regressions between model versions.
Frequently asked questions
Adoption metrics count access: seats provisioned, licences activated, courses completed. Absorption metrics count change: workflow steps redesigned, cycle time reduced, error rate moved. Adoption tells you a tool was available; absorption tells you the work is different because of it. MIT's 95% failure finding describes organisations with high adoption and near-zero absorption.
Capture cycle time per instance, rework rate and escalation rate for one to two weeks from systems that already log timestamps and outcomes, before any tool changes the process. You rarely need new instrumentation; ticketing systems, document stores and approval tools already hold the data. The discipline is to record it before you intervene, not to reconstruct it from memory afterwards.
Only when it is measured net of downstream rework and captured against a pre-baseline. Time saved on a first draft that a human then spends correcting is not saved. Reported as a gross figure from self-report, time saved is one of the least reliable numbers in the whole field, because people consistently feel faster than they are measured to be.
Frequently enough to catch a regression and rarely enough to accumulate signal, typically weekly for an active workflow, against the fixed pre-baseline. What matters more than cadence is that the same metric is reported every time, including the periods when the number is bad. A metric that only appears when it flatters the project is not a metric, it is a highlight reel.
They predict that someone finished a course, and almost nothing about whether a workflow changed. Completion is an adoption metric. It counts an activity, not an outcome. A department can hold a wall of certificates and identical cycle times. If you are reporting completions to a sponsor as evidence of impact, you are reporting the thing most likely to be flat where it matters.
Ready to install the workflow?
Book a free AI Reality Check and build one real thing from your own work, live.