Adoption
Ten questions to ask any AI consulting vendor
Ask a vendor to name the workflow before the engagement starts, to say what they refuse to automate, and to show the harness they leave behind. Including the questions we would fail.


Ask any AI vendor to name the workflow before the engagement starts, to state which workflows they would refuse to automate, and to show the evaluation harness they leave behind. Gartner estimates only around 130 of the thousands of self-described agentic AI vendors are genuine, describing the rest as 'agent washing'.
- Gartner: only about 130 of thousands of self-described agentic AI vendors are genuine: the rest is 'agent washing'.
- A real vendor names the target workflow before the contract, not after.
- A real vendor can tell you what they would refuse to automate and why.
- MIT found buyers of external tools outperformed internal builders.
- Ask the questions the vendor would struggle to answer, including how many engagements they have shipped.
Ask any AI vendor to name the workflow before the engagement starts, to state which workflows they would refuse to automate, and to show the evaluation harness they leave behind. Gartner estimates only around 130 of the thousands of self-described agentic AI vendors are genuine, describing the rest as 'agent washing'.
Key takeaways
- Gartner: only about 130 of thousands of self-described agentic AI vendors are genuine.
- A real vendor names the target workflow before the contract, not after.
- A real vendor can tell you what they would refuse to automate, and why.
- Ask the questions the vendor would struggle to answer, including us.
Why most vendor evaluations ask the wrong things
Most AI vendor evaluations test the wrong surface. They ask about model choice, cloud partnerships, and case studies with impressive percentages, and they come away reassured by answers that predict nothing about whether the engagement will work. The questions that actually predict success are narrower and more uncomfortable, and a vendor's willingness to answer them is itself the strongest signal you will get.
The reason this matters now is the sheer volume of relabelling in the market. Gartner estimates that only around 130 of the thousands of vendors claiming agentic AI are genuine, describing the remainder as 'agent washing': existing products with a new word bolted on (Gartner, 25 June 2025). Ten good questions are how you find the 130.
The ten questions
- Can you name the workflow we would work on before we sign? A real engagement targets a named process. "We'll discover it together" after the contract is a way to bill discovery.
- Which workflows would you refuse to automate, and why? A vendor who cannot answer is selling the thing that fails 95% of the time. This is the whole subject of which workflows you should not automate.
- What baseline will you measure before you build, and who captures it? No pre-baseline, no defensible benefit.
- What evaluation harness will you leave behind? The evaluation harness is the artefact that lets you trust the system after they leave.
- Who owns the code and the harness at the end? The answer must be you, in writing.
- What are the kill criteria? A vendor who has never recommended killing a project has never been honest with a client.
- How is billing tied to deliverables? Milestones against inspectable artefacts, not a lump against a vague scope.
- What governance and audit trail do you build in? Retrofitting an AI governance framework at the production gate is where projects stall.
- How many engagements like this have you actually shipped? Ask it plainly and watch how the answer is handled.
- When are you the wrong choice for us? The vendor who can answer this is the one worth trusting on the other nine.
What a good answer sounds like
A good answer is specific, bounded, and occasionally against the vendor's own interest. To question 2, a good answer names concrete categories ("we would not automate a workflow where the error is expensive and invisible until audit"), rather than "we assess everything case by case." To question 6, a good answer describes a real threshold. To question 10, a good answer names the clients they turn away.
What you are testing across all ten is whether the vendor reasons about your specific situation or recites a capability deck. MIT's NANDA study found that buyers of external tools outperformed organisations that built internally (MIT NANDA, July 2025), but that advantage only holds for the external partners who bring judgement, and judgement is exactly what these ten questions surface.
Red flags: agent washing and the pilot-forever model
Two patterns should end the conversation.
Agent washing. If pressed on question 4, the vendor cannot describe how the 'agent' is evaluated, or the system turns out to be a scripted chatbot with an AI label. The tell is that 'agent' does the work the old product did, with no planning, no tool use, and nothing to evaluate.
The pilot-forever model. The engagement has no exit criteria (question 6) and no production date. Pilots that never end are how a vendor bills indefinitely while the client waits for a value that was never defined. Gartner names unclear business value among the causes of the 40%-plus agentic cancellations it forecasts by 2027; the pilot-forever model manufactures that exact condition.
Questions to ask us
We would fail some of our own questions today, and we would rather tell you that than have you find out.
On question 9 (how many engagements like this have you shipped) our honest answer is: as a new practice, we have shipped zero completed client engagements. The engagement scenarios on this site are illustrative method, not delivered work, and we label them as such. What is real is the domain experience behind them (production EDI and retrieval work), and the method itself, which you can hold us to line by line.
On question 10 (when are we the wrong choice). We are wrong for you if you need a global master-vendor agreement across many countries, if you need indemnity a small practice cannot carry, or if you are a large Indian IT services firm that already trains at a scale we cannot match. Those are on our own refusal list on the about page. A vendor that publishes the clients it turns away is the one telling the truth on the other nine questions, and asking a vendor to disqualify itself is the fastest honesty test there is.
What this means for a GCC transformation owner
Procurement will hand you a vendor scorecard weighted toward the wrong things: certifications, cloud partnerships, headcount. Overlay these ten questions on it. They cost nothing, they take one meeting, and they route past the entire agent-washing layer to the small number of vendors who can reason about your workflow.
And apply question 10 hardest of all. The vendor eager to take every engagement is the one that will keep you in a pilot forever. The one that tells you when to hire someone else, or to do it in-house, is the one to keep on the phone. Before any of this, a scoped adoption diagnostic is how you arrive at the vendor conversation already knowing which workflow you are buying.
Sources
- Gartner, Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, 25 June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- MIT NANDA, The GenAI Divide: State of AI in Business 2025, July 2025. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
Related reading: which workflows you should not automate · what an evaluation harness is · what AI adoption actually costs in India
Read next
- The workflows you should not automateNaming what to leave alone is the first deliverable, not a disclaimer. A vendor who cannot tell you which workflows to keep human is selling you the thing that fails 95% of the time, and the refusal list is the most useful page we can hand a transformation owner.
- Evaluation harnessAn agent evaluation harness is a repeatable test suite that scores an AI agent's outputs against fixed, versioned cases before and after every change, so teams can tell regression from variance.
- AI governance frameworkAn AI governance framework is the documented set of policies, roles, controls and records that determine who may deploy an AI system, on what data, with what testing, and who is accountable when it fails.
More from the blog
- What AI adoption actually costs in IndiaThe market is quote-gated, so buyers cannot benchmark. What actually drives the number, what each engagement type buys, and our real ladder ranges published in full.
- The pilot-to-production checklist for enterprise AIMost AI pilots never reach production because they never had an owner, a baseline, or an evaluation harness. Eleven checks, and the exit criteria for killing a pilot cleanly.
- How to choose the first workflow to automate (and three you should not)The first workflow decides whether the whole programme survives. A four-axis scoring grid, the case for back-office over customer-facing, and the three types to leave alone.
Frequently asked questions
Agent washing is relabelling existing software (a chatbot, a rules engine, a scripted automation) as an 'AI agent' to ride demand. Gartner estimates only around 130 of the thousands of vendors claiming agentic AI are genuine. The test is not the label. It is whether the system plans, uses tools, and can be evaluated against versioned cases, or whether 'agent' is just the new word for the old product.
It depends on the shape of the work, not on a general preference. A boutique wins when the deliverable is working code in one named workflow and you need the team in the room. A large firm wins when the programme spans many countries or needs indemnity a small firm cannot carry. The honest vendor of either size will tell you when they are the wrong choice.
Ask for a reference you can call, ask what the pre-baseline was and who measured it, and ask what the vendor would have done differently. A case study with an achieved percentage but no baseline and no reachable reference is a marketing artefact. Be especially wary of results with no stated measurement method: an unfalsifiable success claim is not evidence.
The named workflow, the pre-baseline and who captures it, the target delta and the date it is measurable, the evaluation harness as an explicit deliverable, the code-ownership terms, the kill criteria, and milestone billing tied to inspectable artefacts. A statement of work that promises a vague productivity gain across a function, with no baseline and no exit criteria, is a statement of hope.
The code, the evaluation harness, the golden-case set, and the documentation: everything needed to run and improve the system without the vendor. A vendor that keeps the harness or the code creates a dependency that is good for them and bad for you. Insist on ownership in writing before the engagement, because it is far harder to negotiate afterwards.
Ready to install the workflow?
Book a free AI Reality Check and build one real thing from your own work, live.