Skip to content
Chokmah

Adoption

Writing an AI business case that survives CFO review

AI business cases fail because they cannot be falsified. Name one workflow, state a measured pre-baseline, and write the kill criteria in advance. A one-page template.

Abstract slug-seeded mark in the brand blue-to-violet gradient standing in for the editorial lead's byline avatar
RileyEditorial lead: AI adoption practice · 20 August 2026 · 6 min readComposite editorial persona. Articles are written and reviewed by the Chokmah practice team.
Glass dashboard cards illustrating the blog cover on a light background

An AI business case survives CFO review when it names one workflow, states a measured pre-baseline, and defines kill criteria in advance. Vague claims fail because they cannot be falsified: METR found developers believed AI made them 20% faster while measurement showed 19% slower. Baseline first, then commit.

  • A business case that cannot be falsified will be rejected by any competent CFO.
  • Name one workflow, not a portfolio of vague productivity gains.
  • State a measured pre-baseline before the project starts, not a remembered one.
  • METR, 2025: developers measured 19% slower believed they were 20% faster: do not trust self-report.
  • Write kill criteria in advance so stopping is a decision, not an admission.
An AI business case survives CFO review when it names one workflow, states a measured pre-baseline, and defines the kill criteria in advance. Vague productivity claims fail because they cannot be falsified: METR's randomised trial found developers believed AI made them 20% faster while measurement showed 19% slower. Baseline first, then commit.

Key takeaways

  • A business case that cannot be falsified will be rejected by any competent CFO.
  • Name one workflow, not a portfolio of vague productivity gains.
  • State a measured pre-baseline before the project starts.
  • METR, 2025: developers measured 19% slower believed they were 20% faster.

Why do AI business cases get rejected?

Because they cannot be checked. The typical AI business case promises a percentage productivity gain across a function or the whole organisation, and a competent CFO reads that as an unfalsifiable claim: a number that can never be proven wrong because it was never defined precisely enough to test. Unfalsifiable claims get rejected, or worse, approved and then quietly forgotten when nobody can say whether they came true.

The fix is to make the case small enough to be falsifiable. One named workflow, one measured baseline, one target, one date by which the delta is known. That is a claim a CFO can approve because it is a claim they can later audit. The move from a portfolio promise to a single measurable workflow is the entire difference between a case that survives and one that does not, and it maps directly onto workflow absorption. You are committing to a change in the work, not to activity.

Baseline before benefit: what to measure first

You cannot claim a benefit you cannot compare against, and the comparison has to be captured before the project starts, because a baseline reconstructed from memory afterwards is worthless. Memory is generous in a predictable direction.

Measure three numbers on the target workflow for two weeks before any tool touches it: cycle time per instance, rework rate, and escalation rate. Pull them from systems that already log them. This pre-baseline is the spine of the whole case, because every benefit claim you make later is a delta against it. Skip it and your December claim is a measured present against a remembered past, which is not a measurement at all.

How to write kill criteria a CFO will accept

A CFO trusts a business case more, not less, when it states the conditions under which you will stop. Kill criteria are pre-agreed thresholds: below this task-success rate, above this rework rate, past this cost per successful task, the project stops. Writing them in advance does two things: it proves you are not emotionally committed to a sunk cost, and it converts stopping from a personal failure into a rational trigger.

| Direction | Example threshold | What it prevents |
|---|---|---|
| Go-live | Task success ≥ target on the golden set | Shipping something that does not work |
| Stop | Rework rate erases the measured time saved | Running a project that costs more than it saves |
| Stop | Cost per successful task exceeds the workflow's tolerance | Silent cost escalation |

Gartner names escalating costs and unclear business value among the reasons it expects over 40% of agentic AI projects to be cancelled by the end of 2027 (Gartner, 25 June 2025). Kill criteria are how a project gets stopped on time and on purpose, instead of cancelled late and expensively.

Counting time saved without lying about it

The single most common way AI business cases overstate benefit is by counting a faster draft and ignoring the corrections it triggers. Real time saved is cycle-time reduction net of downstream rework. If the AI produces a draft in a third of the time but a human spends the saved time fixing it, the net saving is small and your case should say so.

This is where self-report becomes dangerous, because people genuinely believe the tool sped them up even when it did not. METR's randomised controlled trial is the cleanest evidence available: experienced developers were measured as 19% slower using early-2025 AI tools on real tasks in mature codebases, while estimating afterwards that the tools had made them about 20% faster (METR, July 2025). METR has since noted the result is historical and does not necessarily reflect current tools or workflows (METR, February 2026). The tools have improved; the human tendency to feel faster than the measurement has not. Build your case on the measurement, and read the METR result in full before you rely on any self-reported gain.

Where the 95% failure statistic helps your case

Counter-intuitively, the strongest thing you can put in an AI business case is the failure rate. MIT's NANDA study found 95% of enterprise GenAI pilots produced no measurable P&L return (MIT NANDA, July 2025). Your CFO already suspects the base rate is grim. Naming it, and then showing exactly what your case does differently (one workflow, a real baseline, an evaluation harness, written kill criteria) reads as candour and competence rather than naivety.

A business case that hides the base rate looks like every failed pilot's business case did. A business case that opens with the base rate and then dismantles it for your specific workflow looks like the 5% that worked.

A one-page template

One page, six lines. If it does not fit, the case is not focused enough.

  1. Workflow: the single named process this case is about.
  2. Baseline: cycle time, rework rate, escalation rate, measured over two weeks starting on a stated date.
  3. Target: the delta on the primary metric and the date it becomes measurable.
  4. Cost: build cost plus running cost per successful task.
  5. Kill criteria: the thresholds, both directions, that end the project.
  6. Owner: the workflow owner who is accountable for all five lines above.

What this means for a GCC transformation owner

You are the person whose name goes on line six. That is uncomfortable and it is also your leverage, because a CFO backs a case with a named, accountable owner and a falsifiable claim far more readily than a diffuse promise from a central team. Make the case small enough that you can stand behind every number in it.

The reflex under pressure is to inflate the case to make it exciting: bigger scope, bigger percentage, softer definitions. That reflex is exactly what produced the 95%. The case that survives review and then survives reality is the narrow, measured, killable one. Write that one.

Sources

  1. METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
  2. METR, We are Changing our Developer Productivity Experiment Design, 24 February 2026. https://metr.org/blog/2026-02-24-uplift-update/
  3. MIT NANDA, The GenAI Divide: State of AI in Business 2025, July 2025. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
  4. Gartner, Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, 25 June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027

Related reading: the six AI metrics that are not vanity · what workflow absorption means · book an adoption diagnostic

Frequently asked questions

It depends entirely on the workflow, so any universal number is a warning sign. The credible version of the question is: on this named workflow, with this measured baseline, what delta over what period clears our hurdle rate? A CFO trusts a specific answer to that far more than a portfolio-level claim of X% productivity across the organisation, which cannot be checked.

Measure cycle time per workflow instance before and after, from system timestamps rather than self-report, and subtract the rework the AI output creates downstream. Time saved that ignores rework is not time saved. Counting a faster draft while ignoring the corrections it triggers is the most common way AI business cases overstate benefit and lose credibility on review.

Yes. Stating openly that MIT found 95% of GenAI pilots return nothing, and then explaining what your case does differently, is more persuasive than pretending the base rate is favourable. A CFO already suspects the base rate is bad. A business case that names it and shows a specific mitigation reads as candid rather than naive, which is what wins the approval.

Kill criteria are the pre-agreed thresholds below which you stop the project: for example, a minimum task-success rate or a maximum rework rate that the workflow cannot exceed. Writing them before the project starts turns stopping into a rational decision rather than a personal failure, which is the only way projects actually get stopped on time instead of running forever.

The workflow owner: the person accountable for the process the case aims to change. Only they can commit to the baseline, the target and the kill criteria and be answerable for them. A business case owned by a central AI team that does not run the workflow has nobody who can be held to its numbers, which is precisely why such cases drift.

Ready to install the workflow?

Book a free AI Reality Check and build one real thing from your own work, live.