Point of view
The workflows you should not automate
Naming what to leave alone is the first deliverable, not a disclaimer. A vendor who cannot tell you which workflows to keep human is selling you the thing that fails 95% of the time, and the refusal list is the most useful page we can hand a transformation owner.

Do not automate workflows where the error is expensive and invisible, where rules change faster than you can re-evaluate, where the process is undocumented tacit judgement, or where accountability cannot be delegated. Naming these first is a diagnostic obligation, not a disclaimer.
- The refusal list is the first deliverable of a diagnostic, not a disclaimer.
- MIT, 2025: 95% of pilots returned nothing measurable; automating the wrong workflow is a leading cause.
- Four leave-alone categories: expensive-invisible errors, fast-changing rules, tacit judgement, undelegable accountability.
- India's AI Governance Guidelines (Nov 2025) keep accountability with named humans, not systems.
- The right answer for a leave-alone workflow is usually a checkpoint or assist, not full automation.
The evidence
95% of enterprise generative AI pilots produced no measurable P&L return; the root cause MIT identifies is organisational, and automating the wrong workflow is a leading form of it.
Over 40% of agentic AI projects will be cancelled by end of 2027, with inadequate risk controls among the three named causes.
India governs AI through existing law under MeitY's AI Governance Guidelines released 5 November 2025, keeping accountability with named human owners rather than delegable systems.
Why is refusing work a diagnostic output?
Our first non-negotiable is that we tell the client which workflows not to automate. If we cannot, we are selling the thing that fails 95% of the time (MIT, July 2025). That sentence is the whole page.
Most vendor engagements treat the refusal list as a disclaimer buried in a statement of work. We treat it as the primary deliverable of a diagnostic, because it is the part that protects you from the expensive failure. Deciding what to automate is easy and everyone will help you. Deciding what to leave alone requires someone willing to shrink their own scope, and that is rare enough to be a differentiator.
There are four categories we will not automate, and the reasons are specific, not squeamish. A workflow lands in one of them because of how its errors behave, how fast its rules change, whether its judgement is written down, or who is legally accountable when it goes wrong. Understanding these is inseparable from designing where the human stays in the loop.
Over 40% of agentic AI projects will be cancelled by end of 2027, with inadequate risk controls among the three named causes. >Gartner press release (2025)
A refusal list is a risk control. Automating a workflow whose errors are expensive and invisible is precisely how a project accrues the incidents that get it cancelled. Saying no to the wrong workflow protects the budget for the right one.
Category one: where the error is expensive and invisible
Do not automate a workflow whose mistakes are costly and do not announce themselves. Automation is safe when errors are cheap, or loud, or both: a mistranscribed note that the next human catches immediately is fine. It is dangerous when an error is expensive and silent, because the system will produce wrong output confidently and nobody will know until the cost has compounded.
A misclassified transaction that flows downstream into a reconciliation three weeks later is an invisible error. A generated summary that drops the one clause that mattered is an invisible error. The demo never shows these, because the demo runs the happy path, and the happy path is exactly where invisible errors do not appear.
The test we apply is simple: if this step produced a wrong answer, how long until a human noticed, and what would it cost by then? Where the answer is "long" and "a lot," the workflow either stays human or gets a mandatory checkpoint: never full automation.
Category two: where the rules change faster than you can re-evaluate
An automated workflow is only as correct as the last time you evaluated it. If the rules governing the work change faster than you can update and re-test the system, the automation is confidently applying yesterday's policy to today's cases, and the drift is silent.
This is where an evaluation harness is not optional but also not sufficient. The harness tells you the system still passes its cases; it does not tell you the cases are still the right ones. A workflow whose ground truth shifts weekly (a fast-moving pricing rule, an evolving regulatory interpretation, a partner requirement that changes by email) will outrun any re-evaluation cadence you can afford, and the maintenance cost quietly exceeds the saving.
We would rather tell you this at the diagnostic than discover it in month four of a sprint.
Category three: undocumented tacit judgement
Some workflows run on judgement that lives only in a person's head. There is no document, no rule set, no labelled history: just an experienced operator who knows this vendor is always late, that this exception is fine and that one is not, and why. When you ask them to explain, they struggle, because the knowledge is tacit.
You cannot automate what you cannot articulate, and you certainly cannot evaluate it. There is no ground truth to test against. The honest sequence here is not "automate," it is "document first, and decide later." Often the act of documenting the tacit rules is itself the valuable project, and it produces the labelled data that would make automation possible in a future engagement. Trying to skip straight to automation on tacit judgement is how a pilot generates plausible, wrong decisions at scale. This is one reason we run a licence-utilisation audit and interviews before proposing any build.
Category four: where accountability cannot be delegated
Some decisions carry accountability that a named human must hold by law or by contract. India governs AI through existing statute under MeitY's AI Governance Guidelines, released 5 November 2025, which keep responsibility with human owners rather than the system (IAPP, 2025). A trading-partner document with legal force, a compliance sign-off, a decision that affects a person's employment or credit: these can be assisted by AI, but the accountable act cannot be delegated to it.
The distinction is between drafting and deciding. An agent can prepare, retrieve, summarise and propose; a human must own the binding decision, on the record. Designing that boundary well is most of the work in a regulated workflow, and it is the difference between a system that speeds a human up and one that quietly moves accountability onto software that cannot hold it.
Our EDI work is the clearest case: we will assist the exception queue all day, and we will not let an agent alter a legally binding trading-partner document without human sign-off. An engagement scenario for EDI exception triage draws that line explicitly.
What does this mean for a transformation owner?
Ask every vendor, including us, to hand you the refusal list before the scope. A vendor who automates everything is not confident; they are unscoped, and Gartner's inadequate-risk-controls failure is where that ends. The refusal list is the fastest test of whether someone has actually looked at your work or is selling a capability.
Where a workflow lands in one of the four categories, the recommendation is rarely "do nothing." It is usually a smaller, safer intervention: a checkpoint instead of full automation, documentation before automation, an assist that leaves the human deciding. Those are less impressive on a slide and far more likely to survive contact with production.
Disclosure: Chokmah is a new practice with no completed client engagements. This page is method and published evidence. But the refusal discipline is the part of the method we are most willing to be judged on, if we ever propose automating a workflow that belongs in one of these four categories, we have failed our own first non-negotiable.
Get the refusal list before the roadmap
A two-week diagnostic names the three workflows to automate and the ones to leave alone, with the reasoning, in writing.
Book an Adoption Diagnostic · How to pick the first workflow
Frequently asked questions
When should you not use AI?
Avoid full automation where the error is expensive and invisible, where the rules change faster than you can re-evaluate, where the process runs on undocumented tacit judgement, or where legal accountability cannot be delegated. In those cases the right move is usually a human checkpoint or an assist, not a hands-off agent.
What is an invisible error?
An invisible error is a wrong output that does not announce itself and is caught only later, downstream, once the cost has compounded: a misclassified transaction, a summary that drops the one clause that mattered. Automation is dangerous exactly where errors are both expensive and silent, because the demo only ever shows the happy path.
Can you automate a compliance workflow?
You can assist one (retrieval, drafting, summarisation, proposal), but the binding decision and its accountability must stay with a named human. India's AI Governance Guidelines keep responsibility with human owners. The design work is drawing the line between what the agent prepares and what the person decides on the record.
How do you document tacit knowledge before automating?
By shadowing the experienced operator, capturing the exceptions and the reasons behind them, and turning that into an explicit rule set and a labelled history. Often the documentation is the valuable project in itself, and it produces the ground truth that any future automation would need to be evaluated against.
Do you ever recommend against an engagement?
Yes. If a diagnostic finds that the candidate workflows all fall into the four leave-alone categories, or that there is no baseline to measure against, the honest recommendation is to not build yet. Telling a client which workflows to keep human is our first non-negotiable, and sometimes the answer is all of them for now.
Key terms
- Human in the loopHuman in the loop is a workflow design in which a person reviews, approves or corrects an AI system's output at defined checkpoints before it takes effect, keeping accountability with a human.
- Evaluation harnessAn agent evaluation harness is a repeatable test suite that scores an AI agent's outputs against fixed, versioned cases before and after every change, so teams can tell regression from variance.
- AI governance frameworkAn AI governance framework is the documented set of policies, roles, controls and records that determine who may deploy an AI system, on what data, with what testing, and who is accountable when it fails.
More points of view
- Why 95% of GenAI pilots fail, and what the surviving 5% did differentlyThe 95% figure is contested and imperfect, and it still describes your pilot. The failure is not model quality. It is that nobody instrumented the workflow the tool was supposed to change, so no result could ever have been measured.
- Gartner says 40% of agentic AI projects will be cancelled. Here is which 40%.The 40% cancellation rate is not bad luck or immature technology. It is three named, predictable failures (escalating cost, unclear value, absent risk controls) every one of which is decided before the contract is signed, by whether anyone named the workflow first.
- Measure the workflow, not the loginsLicence logins, completions and prompt counts measure that a tool was opened, not that the work changed. The 95% failure rate is a measurement failure as much as an execution one: organisations counted adoption and never instrumented absorption.
Frequently asked questions
Avoid full automation where the error is expensive and invisible, where the rules change faster than you can re-evaluate, where the process runs on undocumented tacit judgement, or where legal accountability cannot be delegated. In those cases the right move is usually a human checkpoint or an assist, not a hands-off agent.
An invisible error is a wrong output that does not announce itself and is caught only later, downstream, once the cost has compounded: a misclassified transaction, a summary that drops the one clause that mattered. Automation is dangerous exactly where errors are both expensive and silent, because the demo only ever shows the happy path.
You can assist one (retrieval, drafting, summarisation, proposal), but the binding decision and its accountability must stay with a named human. India's AI Governance Guidelines keep responsibility with human owners. The design work is drawing the line between what the agent prepares and what the person decides on the record.
By shadowing the experienced operator, capturing the exceptions and the reasons behind them, and turning that into an explicit rule set and a labelled history. Often the documentation is the valuable project in itself, and it produces the ground truth that any future automation would need to be evaluated against.
Yes. If a diagnostic finds that the candidate workflows all fall into the four leave-alone categories, or that there is no baseline to measure against, the honest recommendation is to not build yet. Telling a client which workflows to keep human is our first non-negotiable, and sometimes the answer is all of them for now.
Bring the evidence to your team
We walk in with the failure rates, then the method. Book a free AI Reality Check.