The 95 percent AI failure figure, and what it counted
Since the middle of 2025, one figure has done more work in public argument about AI than any analysis sitting behind it. The claim, as it usually gets repeated: ninety-five percent of corporate AI pilots produce no measurable return.
The source is The GenAI Divide: State of AI in Business 2025, written by Aditya Challapally, Chris Pease, Ramesh Raskar and Pradyumna Chari, and produced in collaboration with Project NANDA out of MIT. It is worth noting that this is not an institutional MIT publication, which is the first thing the shorthand loses. The evidence base was 52 structured interviews, 153 survey responses from senior leaders at four industry conferences, and an analysis of more than 300 publicly disclosed AI initiatives, gathered between January and June 2025. The report’s own sentence is that 95 percent of organizations are getting zero return. A separate finding, stated on its own terms, is that only 5 percent of custom enterprise AI tools reach production.
So the headline number counts organizations. The version in circulation counts pilots, which the report does not claim.
None of that makes the work worthless, and the direction it points in is probably right. The report’s premise, that enterprise AI spending through 2025 had run ahead of demonstrated return, is not seriously disputed, and the attempt to put a size on the gap was a serious one. The problem is what the number became after publication, which is a verdict carrying far more weight than its method can hold. Six months of data collection, at a single point in time rather than tracked, hardened into a judgment on the technology. A convenience sample of initiatives that happened to have been disclosed publicly got quoted as a population statistic. The underlying dataset was never released and the report was not peer reviewed, which prompted Kevin Werbach of Wharton to argue that if Project NANDA stands behind the claims it should publish the full supporting data, and withdraw the report if it will not.
The other numbers
S&P Global Market Intelligence surveyed more than 1,000 respondents in North America and Europe and reported in March 2025 that 42 percent of businesses had scrapped most of their AI initiatives, against 17 percent the year before, with 46 percent of proofs of concept terminated before reaching production. Gartner forecast in July 2024 that 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025, and has since predicted that 40 percent of agentic AI projects will be canceled by the end of 2027.
Most of the disagreement between those figures is definitional. Scrapping an initiative, abandoning a proof of concept, failing to reach production, being canceled, and showing no effect on the P&L are five different events, measured over different windows. The same company could count as a success on Gartner’s test and a failure on NANDA’s in the same quarter without anything about it changing.
That leaves the headline figures, 30 through 95, each defensible and none of them comparable. For a number quoted in support of starting or stopping major programs, that is an awkward foundation, and it sits upstream of the separate question of whether the pilot was ever designed to become a rollout.
What the number is for
A statistic nobody can check supports whichever conclusion the person quoting it already holds. A vendor uses 95 percent to argue that your current approach is the problem and their platform is the fix. A skeptic quotes the same figure to argue the category is overheated and the budget belongs elsewhere. Neither position has to defend the sampling, because the study is not really what is being cited.
Capital appraisal has worked this way for far longer than AI has existed. Bent Flyvbjerg, writing in the Project Management Journal in 2014 on a dataset spanning roughly seventy years of comparable megaproject records, found that nine out of ten overrun on cost. His explanation is a selection effect. Appraisal favors the applications whose costs have been most understated and whose benefits have been most overstated, which is why, in his formulation, the projects that look best on paper turn out worst in reality. Figures nobody can audit are the raw material for that.
The 70 percent change failure rate shows how durable the raw material is. It traces to Hammer and Champy’s Reengineering the Corporation in 1993, where it appears as a self-described unscientific estimate that 50 to 70 percent of reengineering efforts do not achieve the results intended. Hammer disowned it two years later, writing that there is no inherent success or failure rate for reengineering. It spread anyway, helped by a 2000 Harvard Business Review article by Beer and Nohria that presented it as a brutal fact with no footnote, and it is still being quoted more than thirty years after the estimate its author withdrew.
Findings with better pedigree go the same way. The 2010 Science paper by Woolley, Chabris, Pentland, Hashmi and Malone reported a general collective intelligence factor in groups, correlating with equality of conversational turn-taking and with the proportion of women in the group, and only weakly with members’ individual IQ. That is the version of the finding that entered general circulation in writing on team composition. A 2016 re-test by Bates and Gupta in Intelligence, across three studies, found individual IQ accounted for most of the variation between groups, and neither the turn-taking result nor the gender composition result held. Repetition is not retesting, which is why a framework is worth no more than the study and method a reader can check behind it, and why the ones collected on our Insights page carry both.
The second finding
The NANDA report contains a result worth more to a company deciding what to do next than the headline it is quoted for. Pilots built through strategic partnerships, meaning tools sourced externally, were reported as twice as likely to reach full deployment as pilots built internally.
That deserves the same scrutiny the headline gets, and it does not entirely survive it. The two percentages behind the claim, 66 percent and 33 percent, are described as shares of successful deployments rather than success rates, and two shares of the same total summing to ninety-nine is not evidence that one route succeeds twice as often as the other. It says nothing about how many of each kind were attempted. The build-versus-buy conclusion could still be sound. The arithmetic offered for it does not establish that.
The reception is the more telling part. A finding with a decision attached to it, one a reader could argue with and act on, drew very little of the attention the headline got, and its arithmetic went unexamined in the coverage. The figure that spread was the one that implied no decision at all.
A board can leave the question of whether 95 is the right number to other people and produce its own instead. Counting the pilots started in the last two years, and how many of them are running in production today, is a bounded exercise, provided the window and the test get written down before anyone looks at the answer. That count answers the question the 95 percent figure is usually brought in to settle.
← All field notes