Short answer: AI pilots do not fail for one reason, and treating them as a single category is why the same mistakes repeat. There are five distinct failure types the correct kill, the orphan, the ungovernable, the economics inversion, and the unmeasurable and each has a different cause, a different warning sign, and a different remedy. One of the five is not a failure at all. The most common, and the most expensive, is the last: a pilot that produced value nobody can prove, because no baseline was taken before it started.
First, what the 95% statistic actually measured
Almost every article on this topic opens with the same number, and almost every one of them states it incorrectly. Since it is the number your board has probably seen, it is worth being precise.
The GenAI Divide: State of AI in Business 2025, a July 2025 paper from Project NANDA, a research project affiliated with the MIT Media Lab. What it says, verbatim, is that 95% of the organizations studied were getting zero return that is, no measurable P&L impact from generative AI. Its sample was 52 organizations interviewed, 153 senior leaders surveyed, and a review of 300+ publicly disclosed initiatives.
Three things have been done to that finding on its way into general circulation:
The unit was swapped. “95% of organizations report no measurable return” became “95% of pilots fail.” Those are different denominators measuring different things.
The definition was swapped. The paper defines success as a deployment users or executives described as producing “a marked and sustained productivity and/or P&L impact,” measured six months after the pilot. Falling short of that is not the same as breaking, being cancelled, or delivering nothing.
The scope was swapped. The 5% figure applies specifically to custom, task-specific generative AI tools. In the same chart, general-purpose deployments show a 40% successful-implementation rate and roughly 83% pilot-to-implementation. That half of the finding is almost never quoted.
Two further points of context. The paper is a preliminary working paper version 0.1 on its own cover carrying a disclaimer that its views “do not reflect the positions of any affiliated employers,” so citing it as “MIT research” overstates it. And it has drawn substantive criticism, including from Wharton’s Kevin Werbach, who read it repeatedly and could not locate support for the headline claim.
None of this means AI pilots succeed at a comfortable rate. Better-constructed evidence points the same direction: Gartner reported in January 2026 that at least 50% of generative AI projects were abandoned after proof of concept, citing poor data quality, inadequate risk controls, escalating costs and unclear business value. BCG, surveying more than 1,250 companies, found only 5% achieving AI value at scale. Those numbers hold up. The 95%-of-pilots version does not, and repeating it makes the rest of your argument easier to dismiss.
The more useful question is not what percentage fail. It is how they fail, because the five ways are not interchangeable.
Type 1 The correct kill
What it looks like: The pilot ran, tested a real assumption, and the assumption turned out to be false. The team stopped.
This is counted as a failure in every survey and in most internal reporting. It is not one. It is the pilot working exactly as intended buying a decision cheaply before it became expensive.
A pilot exists to answer a question: can this model reach the accuracy this workflow needs, on our actual data, at a cost that clears our threshold? “No” is a valid, valuable answer. A programme that has never killed a pilot is not a programme with a good record. It is a programme that is not testing anything falsifiable.
The warning sign that you cannot tell the difference: you have no written statement of what the pilot was testing. Without that, every stop looks like a defeat, and organisational memory records “we tried AI and it didn’t work” instead of “we established that this workflow needs 99% and the approach gets 94%.”
The remedy is one page, written before the build. What question is this pilot answering. What result means proceed. What result means stop. Who signs either decision. Teams that write this treat a stop as a completed experiment. Teams that do not treat it as a funeral, and the next proposal is harder to fund.
Type 2 The orphan
What it looks like: It works. Everyone agrees it works. It has been “about to go live” for four months.
The orphan is the most common failure in large organisations and the least dramatic. Nothing broke. The pilot demonstrated value, the demo went well, and then the delivery team’s mandate ended and no one in the business took the thing over.
The mechanics are always the same. A pilot is funded as a project with a defined end. Production is an operating commitment with no end someone owns the outcome, someone owns the system, someone is on call, and someone holds a budget line for the run cost. If that transfer is not designed, it does not happen by goodwill.
The warning sign: ask who owns this in production and get a department name instead of a person’s name. “IT” is not an owner. “The AI CoE” is not an owner.
The remedy: name three roles at kickoff, not at handover the business owner accountable for the outcome, the engineer accountable for the system, and the executive who releases the operating budget. Make the funding transfer part of the pilot’s exit criteria rather than a separate conversation afterwards.
Type 3 The ungovernable
What it looks like: It works, and it cannot pass security, legal or compliance review.
This one is almost always a sequencing failure rather than a technical one. The pilot was built in a sandbox on sample data, with a permissive setup nobody intended to keep. Then it met the production environment, and the questions started: which identity does this run as, what can it access, where does the data go, who is the sub processor, what is the audit record, how do we satisfy change-management controls when a non-human actor makes the change.
Retrofitting answers to those questions is not a configuration exercise. Access redesign, entitlement-aware retrieval and audit instrumentation are architecture, and doing them after the fact routinely costs more than the original build.
The schedule impact is the part teams consistently underestimate. Security review for anything that installs infrastructure in an enterprise environment commonly runs six to twelve weeks. Adding a new model provider is a procurement event a new data processing agreement, subprocessor notification and legal review not a settings change. Neither of these is a reason not to proceed. Both are reasons to start them in week one, in parallel with the build, rather than discovering them in week ten.
The warning sign: security and legal have not been in a meeting yet.
The remedy: invite them to the kickoff and give them a swim lane on the plan with a named owner. The controls that make an agent governable scoped identity, least privilege, entitlement-aware retrieval, full trace logging, defined human approval points are the same controls that make it reliable. They are covered in full in our guide to getting agents into production.
Type 4 The economics inversion
What it looks like: Excellent at pilot volume. Unaffordable at production volume.
A pilot at 200 transactions a day tells you almost nothing about the same system at 200,000. Three costs scale differently and one of them scales badly.
Model and retrieval cost scale roughly linearly and get cheaper over time. Infrastructure scales sub-linearly. Human review cost scales linearly and does not get cheaper and in most pilots it is invisible, because the review was done informally by the project team as part of the build.
That is the inversion. A workflow that routes 15% of cases to a human is fine when 15% means thirty items a day. At production volume it means thirty thousand, and you have accidentally designed a staffing plan instead of an automation.
The second, subtler version: pilots are measured on median cost. Production is paid for at the tail. A run that retries, re-plans and re-reads its context can cost many times the median, and the tail is where the budget actually goes.
The warning sign: the business case quotes a cost per transaction and nobody can say what the exception rate is, or who reviews exceptions at ten times the volume.
The remedy: model cost per completed task at target volume before the build, including retries at the 95th percentile and human review minutes at their loaded rate. Then design the review gate so human effort scales sub-linearly staged automated evaluation handling the confident majority, human judgment reserved for the uncertain and the consequential. Getting that routing right is what makes the economics survive success rather than break on it.
Type 5 The unmeasurable
What it looks like: It shipped. People use it. Nobody can prove it was worth anything, so the funding is not renewed.
This is the most expensive of the five, because the value may well have been real. It simply cannot be demonstrated, because nothing was measured before the pilot started.
Consider what “no measurable P&L impact within six months” actually requires to establish. You need the prior cycle time, the prior error rate, the prior cost per transaction and the prior volume. Most organisations do not have those numbers for the process they are automating. They have an impression. When the pilot ends and someone asks what it delivered, the honest answer is “we think it helped,” which does not survive a budget review.
It is worth sitting with the possibility that a meaningful share of the 95% in that NANDA paper is this value that occurred and was never captured, because the instrument to capture it was never installed.
The warning sign: you cannot state today’s numbers for the process you are about to automate.
The remedy is four measurements, taken in week one, before any build: current cycle time per unit, current error or rework rate, current fully loaded cost per unit, and current volume with its variance. Agree with the business owner which of these the pilot is expected to move and by how much. This is a week of work and it converts an unprovable outcome into a defensible one in either direction.
What the pilots that survive do differently
Across our own eleven production engagements, and consistent with what the NANDA paper’s own findings point to, four things separate the ones that reach production from the ones that stall.
They scope to a workflow, not a capability. “Deploy AI in customer service” has no exit criterion. “Reduce first-response time on tier-one queries in this channel” does.
They start from the back office. The paper’s own strongest signal is that value concentrated in unglamorous operational work process elimination, vendor spend reduction, administrative throughput while the majority of budget went to front-office projects with weaker returns. Boring workflows have clean baselines and measurable outcomes. That is not a coincidence; it is the reason they work.
They treat vendors as service providers, not as technology. Benchmark on operational outcomes on your data, not on model leaderboards. Contract on outcomes with defined exit criteria. A vendor benchmark score is a reason to shortlist and never a reason to buy.
They plan for production from week one. The controls that make a system governable are not a phase after the pilot. They are the pilot. Where the source and target sit in the same organisation, this is why a first production workflow takes six to twelve weeks rather than six to twelve months and it is worth noting the NANDA paper found mid-market organisations moving from pilot to full implementation in roughly 90 days, against nine months or more at large enterprises. Smaller scope and shorter approval chains are an advantage, not a limitation.
Our week-by-week roadmap sets out what that looks like in practice, including the security and procurement tracks that run in parallel.
What a failed pilot actually costs
Three costs, and the first is the smallest.
The build cost is sunk and usually modest a scoped pilot is a small number of engineering weeks.
The opportunity cost is larger: the same team, the same quarter, spent on a workflow with a clean baseline and a defined owner.
The credibility cost is the one that compounds. Every stalled pilot raises the evidentiary bar for the next proposal. The third one has to clear a higher hurdle than the first, and it is judged by people who now describe themselves as having “tried AI.” This is precisely why distinguishing a Type 1 correct kill from a Type 2 orphan matters organisationally. One of those is a completed experiment you can point to. The other is a story about AI not working.
Frequently asked questions
Is it true that 95% of AI pilots fail?
No, not as usually stated.
Who published the 95% AI failure statistic?
Project NANDA, a research project affiliated with the MIT Media Lab.
What is the main reason AI pilots fail?
There is no single reason. Failures fall into five types with different remedies: the correct kill, the orphan, the ungovernable, the economics inversion, and the unmeasurable.
How long should an AI pilot run?
Long enough to answer its written question and no longer.
When should we kill an AI pilot?
When the written exit criterion is not met, and immediately.
How do we measure whether an AI pilot succeeded?
Against a baseline taken before it started: cycle time per unit, error or rework rate, fully loaded cost per unit, and volume with its variance.
Why do enterprises struggle with AI more than smaller companies?
Approval chains and scope.
What is pilot purgatory?
A pilot that neither ships nor stops demonstrated, praised, and permanently pending.
