The Failure Rate Has Not Moved in Twenty-Five Years. That Tells You Where to Look.

Posted on September 8, 2026

0


Last week Gartner reported that 22 percent of organizations have successfully scaled AI across multiple business units. McKinsey reported that 88 percent of organizations now use AI and 39 percent see measurable impact on the bottom line. BCG found 5 percent generating significant value.

Read those as an AI story and you will miss what they are.

The same number, five times

Above Graphic Details

EraThe finding, afterwardNon-success rate
CRM (early 2000s)Gartner reported a 50 percent CRM failure rate; Forrester 47 percent; other analyst reports ranged 30 to 70 percent~50%
Big data (2015–2017)Gartner estimated 60 percent of big data projects fail. In 2017 a Gartner analyst said 60 was too conservative and the real figure was closer to 85 percent60–85%
Digital transformation (2016–2024)BCG, across more than 850 companies, found a 35 percent success rate~65%
Data and analytics governance (2024)Gartner projected that by 2027, 80 percent of governance initiatives will fail80%
AI, scaled (2026)Gartner: 22 percent have scaled across business units78%
AI, measurable impact (2026)McKinsey: 39 percent report bottom-line impact61%
AI, significant value (2026)BCG: 5 percent generating significant value95%

⚠ These are not commensurable measurements. Each firm counts something different — abandoned, over budget, not scaled, no measurable return. The band is real; the series is not. What holds across all of them is that the number never lands anywhere good, and never has.

Five technology generations. Four advisory firms. Twenty-five years of different vocabulary. Nothing that could be called improvement.

This has been visible for forty years, and to better observers than me

I want to be careful here, because there is a version of this argument that claims novelty it has not earned.

Robert Solow observed in 1987 that you could see the computer age everywhere except in the productivity statistics. Paul Strassmann spent the 1990s measuring IT spending against corporate profitability and found no reliable correlation. The Standish Group has documented software project failure rates since 1994. Erik Brynjolfsson has worked the productivity paradox since 1993. Jackie Fenn built the Hype Cycle at Gartner in 1995, which named the disappointment phase and gave it a shape.

And in May 2003, Nicholas Carr published IT Doesn’t Matter in the Harvard Business Review, arguing that information technology was becoming an infrastructural commodity and was therefore a diminishing source of competitive differentiation.

The anomaly is not a discovery. It is one of the best-documented findings in modern business, and it has been ignored for four decades by people who were not being careless.

What the record above does not contain is the detail.

Those findings are aggregates, and aggregates are assembled after the outcome is known. A survey asks how many initiatives succeeded. A correlation compares spending against returns across an economy. A project census counts how many finished. All of it is true, all of it is durable, and none of it can tell you what happened inside any single implementation, because by the time the number is compiled the particulars have been averaged out of it.

I have been tracking the particulars since 2007, and the archive is dated. Specific organizations, specific implementations, specific performance figures, recorded as they happened rather than reconstructed once the result was in. Where a program was covered a decade ago it is still traceable now, through the acquisitions and the rebrands and the quiet disappearances, because each observation carries the date it was made and none of them has been revised to match what came later.

That includes the ones that worked. I have followed Virginia’s eVA program for roughly two decades, in contact with the people who built and ran it, and it has taught me more than most of the failures — because a failure case tells you what was missing, and only a success tracked the same way over the same span tells you what was present.

Two things make it a control rather than an anecdote. The program has since moved off the platform it was built on, and its performance did not depend on that platform. And the entire founding leadership stratum has turned over — the architects have retired or moved on, and people who were junior when I started covering it now hold senior positions — while the results persisted. That matters, because the most sophisticated objection to any success story is that it was a few exceptional individuals rather than anything transferable. A result that survives complete leadership turnover was never lodged in the leaders. It was built into the operating conditions, and conditions are the thing that can be assessed somewhere else.

Without at least one case like that, you cannot distinguish a condition that was absent where things went wrong from a condition that is absent everywhere. An archive that recorded only failures would have no way to know what it was looking at.

That is a different kind of evidence, not a better grade of insight. A retrospective account can only contain what the answer needed. A contemporaneous record contains what was true at the time, including the calls that turned out wrong — which is how you can tell it was not tidied afterward, and why it can still be interrogated for causes rather than confirming a conclusion already reached.

Why an accurate correction does not change the trajectory

Carr is the useful case, because his argument was substantially right and it changed almost nothing.

The reaction was immediate. Microsoft’s Steve Ballmer called it hogwash at a financial analysts meeting. Bill Gates said they disagreed with all of it. Executives at HP, Intel and IBM felt compelled to respond publicly. Carr’s own description is that he was branded a heretic. The Harvard Business Review published seventeen pages of letters, and that same year its staff voted it the best article the magazine had published.

Carr later identified part of the reason himself: the article went directly at the marketing message the industry was running on — be on the cutting edge or be left behind. The parties best equipped to respond were the parties it targeted.

But individual decisions are the least interesting part of it. The reasons that matter are structural, and they would have operated whoever was in the room.

The prescription was unsellable. Carr told organizations to spend less, follow rather than lead, and make IT management boring. No career advances by proposing restraint. No budget line funds it. There is no procurement category called do less than we planned.

It offered no method. Don’t lead is a strategic posture, not something a practitioner can execute on a Monday. Solow, Strassmann, Standish and Brynjolfsson have the same limitation, and it is not a criticism of their work: they were measuring economies, project populations and spending against returns. Aggregate findings cannot tell a specific organization what is wrong with a specific implementation.

Disappointment was already named as a phase. Once the trough of disillusionment is a stage on a curve, a failure rate stops being evidence that the approach is wrong and becomes a sign that you are on schedule.

Failure gets attributed to execution. Poor data, weak governance, insufficient readiness. Each explanation is usually true. Each also leaves the prescription intact for the next era.

And the vocabulary resets. By 2007 the argument had to be relitigated in the language of cloud, then mobile, then digital, then AI. You cannot compare this wave to the last one when none of the words match.

None of this requires anyone to have acted badly. It is the shape of the thing.

What is different here

Not the observation. The altitude.

Every antecedent above worked in aggregates. That is what makes their findings durable and what makes them unactionable. No economy-level correlation has ever told anyone why their implementation is failing.

I work at the level of the single operation, and I trace backward from one failure rather than mapping forward toward an outcome.

In 1998, a defence maintenance operation was running at 51 percent next-day delivery. Every conventional explanation was available — carriers, warehouse throughput, supplier performance, order accuracy. The measures on the board were accurate. The question that mattered was what time of day the technicians submitted their orders, because a next-day requirement is a duration and a duration has a start point. The answer was four o’clock in the afternoon. Procurement cannot act on an order it has not received.

Nobody had established the boundary the performance was being measured from. No aggregate finding would ever have surfaced it, and no scorecard was going to contain it, because time of day was not a performance dimension in that operation or in any other I have seen since.

Delivery moved to 97.3 percent in three months and held for seven years. The platform selection came after the gain was substantially achieved.

That is the difference. Trace backward from a failure and you find the cause. Map forward toward an outcome and you find failures of unknown origin after the fact.

How it gets fixed

Not by arguing that the rate is high. Everyone above already established that, and it changed nothing.

It gets fixed one operation at a time, by establishing whether the conditions the technology is being introduced into can carry it — before the commitment, not after the post-mortem.

That means answering three questions that no vendor evaluation asks and no scorecard contains. What determines the outcome here, and is it currently being measured? Where does the process cross a boundary that nobody owns? And what would have to be true for this to work, independent of which product is selected?

It is a small, unglamorous piece of work. It has no category, no budget line and no natural buyer, which is exactly why the failure rate has held for twenty-five years while everything around it changed.

The number is not a property of any technology. It has survived five of them. It is a property of the conditions the technology arrives into — and conditions can be assessed.

The thing was never the problem

Money gets blamed for a great deal, but the older version of that saying is more careful than the one we repeat. It is not money. It is the pursuit of it, and what gets done with it once it arrives.

Technology is no different. ERP was not the problem. Neither was cloud, or SaaS, and neither is AI. Each arrived with real capability, and in five consecutive generations the capability was rarely what failed.

What repeats is the pursuit. Capability first, conditions later, with the assessment scheduled after the commitment rather than before it. That sequence has now outlived every technology it has been applied to, which is the strongest evidence available that the sequence is the variable.

Reverse it, and the technology stops being the thing that has to work.


Truth Is Believing. Accuracy Is Knowing. Outcome Is Proof.™

-30-

Posted in: Commentary