Skip to content
Chokmah

Point of view

Measure the workflow, not the logins

Licence logins, completions and prompt counts measure that a tool was opened, not that the work changed. The 95% failure rate is a measurement failure as much as an execution one: organisations counted adoption and never instrumented absorption.

Abstract slug-seeded mark in the brand blue-to-violet gradient standing in for the editorial lead's byline avatar
RileyEditorial lead: AI adoption practice · 25 July 2026 · 6 min readComposite editorial persona. Articles are written and reviewed by the Chokmah practice team.

Licence logins measure that a tool was opened; workflow throughput measures whether the work changed. MIT's finding that 95% of enterprise GenAI pilots delivered no measurable P&L impact is a measurement failure as much as an execution one: organisations counted adoption and never instrumented absorption or redesign.

  • Adoption counts access; absorption counts change. Only absorption reaches the P&L.
  • MIT, 2025: 95% of GenAI pilots showed no measurable P&L impact: many were never baselined.
  • METR: developers measured 19% slower while feeling faster: self-report is not a metric.
  • Sign a measurement contract before building: four workflow numbers, a baseline, and kill criteria.
  • A good absorption metric can come back bad; protect whoever has to report it.

The evidence

95% of enterprise generative AI pilots produced no measurable P&L impact, a gap MIT attributes to organisational rather than technical causes.
MIT NANDA, The GenAI Divide: State of AI in Business 2025 (2025)
In a randomised trial, developers measured 19% slower with AI while believing they were about 20% faster: self-report and measurement pointing opposite ways.
METR, early-2025 developer productivity study (2025)
METR states the 19% result is historical and does not necessarily reflect current tools or workflows.
METR, We are Changing our Developer Productivity Experiment Design (2026)

What is the difference between adoption and absorption?

Adoption counts access: seats provisioned, licences activated, courses completed, prompts sent. Absorption counts change: workflow steps redesigned, cycle time reduced, error rate moved. MIT's finding that 95% of enterprise GenAI pilots produced no measurable P&L impact describes organisations with high adoption and near-zero absorption (MIT, July 2025).

Every failing pilot we have read about had good adoption numbers. People logged in. Certificates were issued. The dashboard was green. And the work was unchanged, because nobody was measuring the work. This is the gap we named workflow absorption, and it is the only number that ever reaches the P&L.

The reason the distinction goes unmade is that adoption is easy to measure and absorption is hard. A platform can report a login automatically; measuring absorption means instrumenting a process before you touch it. So organisations measure what is easy, and then wonder why the easy number rose while nothing changed.

95% of enterprise GenAI pilots produced no measurable P&L impact: a measurement failure as much as an execution one. >MIT NANDA, The GenAI Divide: State of AI in Business 2025 (2025)

If the workflow was never baselined, no result could have been detected regardless of how well the tool performed. Absence of measured impact and absence of measurement are indistinguishable from the dashboard.

Why do dashboards reward the wrong behaviour?

A metric is an instruction. When you report weekly logins to a steering committee, you are instructing the organisation to produce logins, and it will: regardless of whether logins do anything. People are not cynical about this; they are responsive. Show them the number that is watched and they will move it.

The METR trial is the sharpest illustration of why self-report cannot be trusted as the metric. Developers measured 19% slower with AI while believing they were about 20% faster (METR, July 2025; METR notes the result is historical and may not reflect current tools). If the people doing the work cannot feel their own speed, then "the team says it helps" is not evidence, and a dashboard built on satisfaction and usage is measuring the wrong thing confidently.

The fix is not more dashboards. It is fewer, harder numbers: the ones that would embarrass you if they came back flat.

The measurement contract we sign at the start

Before we build anything, we agree on the numbers that will judge the work, and we record them. This is a contract, not a preference, because a metric chosen after the result is not a measurement. It is a justification.

For a given workflow, the contract typically fixes: cycle time per instance, rework or error rate, escalation rate, and cost per successful task. We baseline each one for a period before a single line is written. We also write, in the same document, what result would cause us to recommend stopping. A pre-committed kill criterion is the only defence against the perception gap the METR trial exposed.

This is deliberately uncomfortable. It means we can be shown to have failed, on numbers we agreed to in advance, which is precisely why it is trustworthy. A two-week Adoption Diagnostic produces this contract as its first artefact, and an evaluation harness keeps the post-build numbers honest run after run. An engagement scenario for vendor invoice reconciliation shows the shape of a baselined workflow, as method rather than a client result.

What happens when the number comes back bad?

We tell you. This is the part of the contract that most vendor engagements quietly lack, and it is the part that makes the good numbers mean anything.

A workflow that did not improve is not a hidden result to be reframed in the readout. It is a finding, and often a valuable one. It tells you that this workflow was the wrong first candidate, or that the process needs redesign before automation, or that the exception path is larger than anyone thought. Reported honestly and early, a flat number saves the far larger spend of scaling a thing that does not work. That is a materially better outcome than a renewed licence and a satisfied survey.

Reporting a failed workflow to your sponsor is a skill, and we would rather coach you through it than help you avoid it. The alternative (a green dashboard over an unchanged process) is exactly the 95% condition.

What does this mean for a GCC transformation owner?

Change what you report upward. Retire logins, completions and prompt counts from the steering deck, not because they are false but because they are answered questions. Replace them with the four workflow numbers, baselined before the build, and the kill criteria agreed in advance.

Then protect the person who has to report a flat number. If reporting bad news is punished, you will get green dashboards over unchanged work, which is the most expensive outcome available. The whole value of absorption metrics is that they can be bad, and a culture that cannot hear a bad absorption number will regenerate the 95% failure no matter which vendor it hires.

Disclosure: Chokmah is a new practice with no completed client engagements. Every number on this site is published research or method; we have no results to report yet, and we would rather say so than imply otherwise.

Instrument the workflow before you scale the tool

A two-week diagnostic delivers the measurement contract (the four numbers, the baseline, and the kill criteria) before anything is built.

Book an Adoption Diagnostic · See how the ladder is sequenced

Frequently asked questions

What is workflow absorption?

Workflow absorption is the degree to which an AI tool has actually changed how a process runs (steps redesigned, cycle time reduced, errors cut) as distinct from adoption, which only counts access such as logins and completions. Absorption is the part that reaches the P&L, and the part almost nobody instruments.

What is a good AI ROI metric?

One measured against a pre-build baseline on the specific workflow: cycle time per instance, rework or error rate, escalation rate, or cost per successful task. A good metric can come back bad. If a metric can only improve or stay flat, it is a vanity metric, not an ROI metric.

Should AI metrics be tied to headcount reduction?

Not by default, and not as the headline. Tying the metric to headcount makes the workforce hostile to instrumentation and corrupts the data you need. Measure the work: throughput, error rate, cycle time. Whatever headcount decisions follow are a separate conversation, made on clean numbers.

How long before AI ROI is measurable?

Sooner than most expect, if you baseline first. With a workflow instrumented before the build, you can read a signal within weeks of go-live because you have a clean before-and-after. Without a baseline, ROI is never cleanly measurable, which is how pilots run for a year and prove nothing.

Who should report AI metrics?

The workflow owner, not the tool vendor and not IT. The person accountable for the process outcome should own the number, because they are the one who can act on it. And they must be able to report a flat number without penalty, or the metric stops being honest.

Frequently asked questions

Workflow absorption is the degree to which an AI tool has actually changed how a process runs (steps redesigned, cycle time reduced, errors cut) as distinct from adoption, which only counts access such as logins and completions. Absorption is the part that reaches the P&L, and the part almost nobody instruments.

One measured against a pre-build baseline on the specific workflow: cycle time per instance, rework or error rate, escalation rate, or cost per successful task. A good metric can come back bad. If a metric can only improve or stay flat, it is a vanity metric, not an ROI metric.

Not by default, and not as the headline. Tying the metric to headcount makes the workforce hostile to instrumentation and corrupts the data you need. Measure the work: throughput, error rate, cycle time. Whatever headcount decisions follow are a separate conversation, made on clean numbers.

Sooner than most expect, if you baseline first. With a workflow instrumented before the build, you can read a signal within weeks of go-live because you have a clean before-and-after. Without a baseline, ROI is never cleanly measurable, which is how pilots run for a year and prove nothing.

The workflow owner, not the tool vendor and not IT. The person accountable for the process outcome should own the number, because they are the one who can act on it. And they must be able to report a flat number without penalty, or the metric stops being honest.

Bring the evidence to your team

We walk in with the failure rates, then the method. Book a free AI Reality Check.