How often does an initiative quietly change what it was trying to achieve — and what does that do to your ROI?
There is a question I went looking for an answer to this week, and could not find one.
Of all the enterprise initiatives that fall short of what they originally promised, what percentage quietly replace the original objectives with new ones — and then report success against the replacements?
I expected the number to be somewhere. It is not. And the absence turns out to be more instructive than the number would have been.
What the research actually covers
The literature on project success is substantial. There are multidimensional models of success. There are benefits realization frameworks. There are maturity models for defining criteria, and there are structured processes for reaching agreement on them with stakeholders.
The shifting of goalposts is acknowledged within that literature — usually as a frustration inflicted on project teams, and usually with the same remedy attached: define your criteria clearly at the outset and get stakeholders to agree.
That is prevention advice. It is not measurement.
Nobody counts how often the goalposts actually move. Nobody tracks what the original business case said against what the program later reported. And the field is candid about its own footing: decades of research have left practitioners and scholars with a vague notion of what project success even is, settling for conflicting attributions of a phenomenon that remains elusive.
So we have a large body of work on how success should be defined, and almost nothing on what happens to those definitions once a program is under way and the original numbers are no longer reachable.
That is a strange place for a discipline to be. It is the equivalent of an extensive literature on how to set a budget, and no data at all on how often budgets get rewritten to match the spend.
Why the original number was never a fair benchmark
Based on the research, and on nearly three decades of contemporaneous observation, it is reasonable to conclude that the following is a contributing cause of the gap.
An original business case is written at the point where you know the least and must promise the most.
You know the least, because it is produced before anyone has traced how the organization actually operates. You must promise the most, because it has to clear a funding gate against competing claims on the same money.
Put those two conditions together and optimism is not a character flaw of the people writing it. It is the function of the document. The business case exists to secure a decision, and it is written by people who do not yet know what they will find.
Which means the original target and the eventual measured reality were never really commensurable. One was a persuasion instrument. The other is a measurement. Treating the gap between them as straightforward failure is a category error — and it goes some way toward explaining why organizations reach for the revision rather than the reckoning. Nobody wants to be measured against a number that was never a measurement, and should never have been treated as one.
The problem is not that targets get revised. The problem is that they get revised silently, and the original stops being referenced.
What that does to ROI
This is where it stops being a philosophical point.
If you are assessing the return on an initiative, you are comparing outcomes against a commitment. If the commitment has been restated — and the restatement is not documented, and the original is no longer retrievable — then your ROI assessment is measuring against a moving target and reporting a fixed number.
You will get an answer. It will look precise. It will mean nothing.
And this failure mode is more durable than overspending. An overrun eventually exhausts its budget and forces a reckoning. A redefined success criterion never runs out of anything. The program succeeds, permanently, against terms nobody has to defend, because the terms it was originally funded against are no longer in the room.
Which brings me to the one instance I can put a date on.
The one I can prove
In 2017 I was covering a public-sector e-procurement implementation that had been struggling for two years. I interviewed the executive who had taken over the account and asked about the adoption figures that had been the subject of my earlier reporting.
The answer was that those numbers were “not a relevant marker of success,” and that the initiative was “already a success” — having reached what was described as a critical mass of onboarding, with the remaining work characterized as optimizing usage.
I want to be careful here, because the easy reading is the wrong one. Nothing about that was evasive. It was said openly, on the record, to a journalist, for publication. She also placed responsibility for the adoption problems with the client rather than the platform — and on that point I agreed with her at the time and still do. A vendor cannot supply readiness an organization has not built, and blaming the platform for what the organization did not do is one of the more common errors in this field.
What interests me is not motive. It is that a metric can be set aside as “not relevant” without anyone lying, and the record afterwards contains no trace of the original standard. That is the mechanism, and it does not require bad faith to operate. It only requires that nobody preserved what was being replaced.
Which is why I wrote down my own test at the end of that piece: greater engagement, supplier involvement, and quantifiable taxpayer benefit — the last of those being the actual measure of success.
That post has not been edited since it was published on 3 August 2017. The publication and modification timestamps are identical, and both are held by a third party, not by me. The same applies to every post in this archive.
Nine years on, none of the three has been publicly demonstrated. The portal endures. Participation is still not universal. The system survived. That is not the same thing as the system having worked.
I do not offer that as a percentage. It is just one case example. But it is one case with the criteria written down in advance, dated by someone other than me, and checkable.
What the gap actually costs
The obvious way to price a revision gap is arithmetic: what was committed, what was delivered, subtract. That is worth doing and it is not the answer.
A business case does not merely describe a benefit. It authorizes a decision. A spend. A vendor. A sequence. A headcount plan. A set of alternatives that were considered and set aside. Every one of those was decided on the strength of a number.
If that number was never achievable, the shortfall in benefit is the smallest part of what it cost you. The real cost is everything that was decided on its strength — and none of it appears on the program’s ledger, because those decisions were booked elsewhere, to other budgets, in other years.
An overrun costs you the overrun. A revised criterion costs you the path you did not take.
And it costs you the next initiative.
People remember. A program declared successful against terms nobody recognizes teaches an organization that success is something announced rather than achieved. So the next time someone stands up to explain why this one is different, they are talking to a room that has already learned otherwise. That resistance gets called change management. Most of it is memory.
So the practical calculation is three questions rather than a formula:
What did this number authorize? Not what did it promise — what did it permit. The spend, the selection, the sequencing, the reorganization.
What else was on the table when it was authorized? The alternatives that lost. Usually documented, usually in the same paper, usually never revisited.
Would that decision have been made against the number we are reporting today?
If the answer to the third question is no, you have your figure. It is the difference between the path taken and the path that would have been taken — and it is nearly always larger than the benefit shortfall, because it compounds across every year since.
Two things follow from that, and neither is comfortable.
The first is that you cannot run this calculation without the original. Not a summary of it, not a revised version — the document as approved, in its own terms. Which is precisely what a silent revision removes. The organizations most exposed to this cost are the ones least able to compute it.
The second is that the calculation gets more expensive the longer you wait, and not linearly. Every year the original stays unexamined is another year of decisions taken on a foundation nobody has checked, each becoming the premise for the next.
And underneath those decisions, something quieter is happening.
When a system does not work the way it was promised, front-line practitioners do not stop working. They route around it. In 2005 I stood in front of a room of industrial distributors in Calgary and put the number on the screen: 79 percent of MRO purchases in Canada were going off-contract. Not because buyers were careless or undisciplined. Because getting the job done required it.
What follows is not resistance. It is concession. The organization tacitly accepts the workaround rather than examining why it became necessary — and the workaround quietly becomes the process. Nobody decides this. Nobody writes it down. There is no meeting.
A few years on, “we have always done it this way” no longer describes the system that was funded. It describes the accommodation that replaced it. And by then the original business case is not merely unexamined. It is describing something that no longer exists.
Which is the argument for doing it now rather than at the next reorganization, when the people who wrote the business case have moved on and the only surviving version of the target is the one that was met.
The fourth name
I wrote recently about a principle I had been applying for nearly three decades before discovering that other people had arrived at it in fields I know nothing about. Carl Jacobi, who inverted mathematical problems that would not yield to the forward formulas. Charlie Munger, who asked not how to succeed but how he would fail, and then stayed away from there.
There is a fourth name, and he is the one who counted.
Bent Flyvbjerg, Professor Emeritus at Oxford’s Saïd Business School and the most-cited scholar in the world in megaproject management, has spent decades assembling a database of more than sixteen thousand projects, spanning twenty fields across a hundred and thirty-six countries. From it he derived what he calls the Iron Law of Megaprojects: over budget, over time, under benefits, over and over again.
His remedy is reference class forecasting, and the logic will be familiar by now. Do not forecast from inside your own plan. Identify a class of comparable completed projects, establish the actual distribution of what happened to them, and calibrate your estimate against that distribution rather than against your own intentions.
Which is to say: do not reason forward from what you hope. Go and look at what actually happened, failures included, and reason from there.
That is one of the significant benefits of the Procurement Insights archive and the Provenance Ledger behind it — nineteen years of papers, videos, podcasts, reports and, of course, articles, dated by third parties and never edited after publication. A project database records what happened. A contemporaneous archive records what was said would happen, when, and by whom — and then what happened. The ledger exists to keep those separate, so that a recollection never quietly becomes a record.
And here are the figures that should stop anyone reading this.
Reporting on those sixteen thousand projects in How Big Things Get Done, Flyvbjerg and his co-author Dan Gardner found that 8.5 percent came in on cost and on time. Half of one percent came in on cost, on time, and on benefits.
Benefits. The dimension almost nobody hits — and the only one of the three with no external referent. A cost overrun is arithmetic; the number is the number. A schedule overrun is a calendar. A benefit shortfall is a matter of definition, and definitions can be revised.
It is worth noting how carefully Flyvbjerg and Gardner had to guard against exactly that. Their overruns are measured against the budget originally approved at the decision-to-build stage — deliberately not against figures re-baselined partway through, because business cases are re-baselined often enough that measuring against the revised number would have hidden the overrun altogether.
They had to build the defense against moved goalposts into the measurement itself, or the dataset would have told them nothing. That is independent corroboration, from sixteen thousand cases, that the practice is common enough to invalidate a study.
The half of one percent and the missing revision statistic are the same finding, approached from two directions.
What that collapse looks like at close range
I went back to something I wrote in 2008 to see whether it explained the gap between those three numbers.
It was a white paper on SAP’s Procurement for Public Sector offering, published through CATA Alliance in Ottawa. I gathered every case reference I could find — the vendor’s own published successes, the integrator case studies, and, deliberately, the failures that neither group was advertising. Roughly two dozen named organizations in total.
It is not a dataset. There is no defined population, so there is no denominator, and the two halves were assembled by opposite methods — the successes came from the vendor’s site, the failures came from going looking for failures. No percentage calculated from it would mean anything, and I am not going to offer one.
But you can see the mechanism working, and it explains the aggregate.
Of the named organizations, two claimed both cost and time. One county completed a plain-vanilla implementation with no customizations in nine months, on time and on budget. One liquor corporation reported twelve months, on time. That is roughly the same neighborhood as 8.5 percent, and at that sample size the resemblance is coincidence rather than corroboration. I note it only because it points in the same direction.
Of the named organizations, not one published what it had originally committed to deliver alongside what it delivered.
Not one. Read the success cases and the pattern is unmistakable. One city was anticipating a full return within five years — a projection, not an outcome. One council would save 72 percent of the purchase and installation cost over its first five years — future tense. One county “successfully standardized and streamlined business processes,” with no figure attached at all. One school district reported saving 50 to 75 percent of back-end processing costs, against a baseline nobody states.
These are not dishonest documents. They are simply not measurements. Every one of them reports activity, or projection, or a percentage floating free of the number it improved on.
Which is the answer to why the benefits figure collapses to half of one percent. The benefits dimension is barely reported anywhere in a form that could be checked. And a target that was never published in a form anyone could check does not need to be moved. It simply never has to be met.
You cannot move a goalpost that was never planted where anyone could see it.
The reference account
There is one case I have never forgotten, and it is the reason I stopped treating vendor case studies as evidence of anything. The white paper recorded half of it. I published the other half two years later, once I understood what I had actually been looking at.
In October 2007 I was engaged by a county government to assess a proposed business process. It was a short piece of work — days, not weeks.
In December 2007, the vendor courting that county put another county government, in a neighboring state, forward as its case reference. That is standard practice, and nobody should be criticized for it — a reference account is how enterprise software is sold. But something about it did not sit right with the people I had been working for, and they asked me to look at it.
In January 2008 — the following month — the reference county announced that after roughly two to three years and thirty-eight million dollars, its initiative had blown up, creating a political crisis.
Four weeks between being a reference account and being a catastrophe.
It would be easy to read that as deception. I do not think it was, and the honest version is more useful.
A program of that size does not collapse in four weeks. The condition existed in December. What happened in January was an announcement — the county said out loud what had been true for some time, and it made news precisely because it had not been said before. Until then, the vendor’s reference list reflected what the client was saying about itself.
Which is what a reference account is: a self-report, carrying a second party’s name. Nobody had verified it, because verification is not part of the mechanism. There is no step in that process where checking is anybody’s job.
That is a more troubling finding than deception would have been. A lie can be caught. This cannot, because nothing went wrong — every party behaved exactly as the process expects, and the output was still worthless.
When I wrote about it again in 2010, I put it in one line, and I have not improved on it since: today’s case reference could be tomorrow’s embarrassment.
That is what a success set composed of self-reported cases with no verifiable benefits data is actually worth. And it is worth remembering when your own program is being assessed on the strength of somebody else’s reference list — or when yours is about to become somebody else’s.
Four operations, one principle
I want to be precise about what is shared and what is not, because a loose version of this claim is worth less than an accurate one.
These are four different operations, and the differences are not cosmetic.
Jacobi restates. He takes a problem that resists a frontal assault and flips it into its inverse form, where the path becomes visible. The work is mathematical and it is done on paper.
Munger avoids. He imagines a failure that has not occurred, enumerates the ways it could arrive, and steers clear of them. The failure is hypothetical by design; the point is never to meet it.
Flyvbjerg counts. He takes failures that have already finished, aggregates them across an entire class, and uses the distribution to correct a forecast before a new project starts. The failures are real but they are other people’s, and they are closed.
I trace. The failure is not hypothetical, not aggregated, and not finished. It is happening right now, inside an organization, with people in it who have positions attached to how it gets explained. I start at the breakage and follow the mechanism backward until it terminates in a cause — which is routinely somewhere nobody was looking, and frequently outside the department that owns the problem.
That last difference is the one that matters in practice, and it is not a claim of rank. It is a claim about conditions. Three of these operations can be performed with the evidence sitting still. Mine has to survive contact with an operating reality that pushes back while you are working on it — where the people who chose the original path are in the room, and where the cause you are tracing toward may be sitting across the table.
Underneath the four operations, one idea: some problems cannot be solved forward.
I did not learn it from any of them. I developed and proved my method in 1998, with funding from the Government of Canada’s Scientific Research & Experimental Development (SR&ED) Program, having never encountered Jacobi’s maxim and knowing nothing of Munger’s work. You cannot operationalize what you have not read.
Which is the most persuasive thing about it. A preference held by four people who share a reading list is a school. A principle that surfaces separately in pure mathematics, in investing, in the statistics of megaprojects, and in enterprise failure diagnosis — in four operations that do not resemble each other — is not a preference. It is closer to a law.
What the backward approach does to ROI assessment
This is the practical payoff, and it is the reason any of the above matters to someone with a live initiative.
An equation-based approach starts from the goal and builds toward it. Its defining property is that it commits you, structurally, to being right about the path you chose. Every subsequent fact sorts into helps me get there or obstacle. You are no longer investigating; you are defending a destination. And when the destination becomes unreachable, the cheapest available move is not to abandon it — it is to redescribe it.
An agent-based approach starts where the thing is failing and traces backward. Because you began at the breakage rather than the goal, you have nothing to defend. You can follow the evidence out of your own department, past your own decisions, to a cause nobody suspected.
Applied to ROI assessment, that produces three things the forward method cannot:
A retrievable original. The trace forces you to establish what was actually committed to, because that is where the trace terminates. You cannot quietly lose a document you are working backward toward.
A measurement against the operating reality rather than the pitch. The trace tells you which aspects of the operating reality were validated and which were only assumed. Return calculated against validated conditions means something. Return calculated against assumed conditions is arithmetic performed on a wish.
A documented delta. This is the deliverable that matters, and almost nobody produces it: what was assumed, what the trace found, what changed, and why. Not a replacement business case. The original and the revision, side by side, with the evidence between them.
A real and documented initiative, mapped two ways. On the left, the process as designed — every arrow pointing at the target, containing only what someone thought to build in. On the right, the trace: it leaves the mapped process at the second step, loops on itself, and terminates in a cause nobody had thought to include. First published in The Map You Design and the Map You Trace, 29 July 2026.
The test
Because targets should sometimes change. An initiative that discovers its founding assumptions were wrong and carries on toward the original number is not disciplined; it is stubborn. Revision can be the correct output of good diagnostic work.
So how do you tell the difference between a legitimate revision and a moved goalpost?
Three conditions, all required:
- The original is preserved — still retrievable, still referenced, not superseded into silence.
- The reasons are documented — what specifically made the original misaligned with the operating reality.
- The revision is derived from a trace, not declared.
The third is the one that does the work. A revision can be announced openly, on the record, for publication — and still be a moved goalpost, if the original is not preserved and the change is asserted rather than derived from evidence. Openness alone does not qualify it. In the 2017 interview above, the new criterion was stated publicly and without evasion. The original was simply set aside as not relevant. Nothing was hidden. The goalpost moved anyway.
Which is why the missing statistic may be missing for a reason. You cannot count what nobody is required to preserve.
If your organization is running an initiative right now, there is one question worth asking this week, and it does not require an assessment, a consultant, or a methodology:
Can someone produce the original business case — and do the measures we currently report against still map to it?
If the answer is yes, you are in better shape than most.
If the answer is difficult to find out, you have just learned something more valuable than any ROI figure your current reporting would have given you.
-30-
Truth Is Believing. Accuracy Is Knowing. Outcome Is Proof.™
Related
The Number Nobody Collects
Posted on August 4, 2026
0
How often does an initiative quietly change what it was trying to achieve — and what does that do to your ROI?
There is a question I went looking for an answer to this week, and could not find one.
Of all the enterprise initiatives that fall short of what they originally promised, what percentage quietly replace the original objectives with new ones — and then report success against the replacements?
I expected the number to be somewhere. It is not. And the absence turns out to be more instructive than the number would have been.
What the research actually covers
The literature on project success is substantial. There are multidimensional models of success. There are benefits realization frameworks. There are maturity models for defining criteria, and there are structured processes for reaching agreement on them with stakeholders.
The shifting of goalposts is acknowledged within that literature — usually as a frustration inflicted on project teams, and usually with the same remedy attached: define your criteria clearly at the outset and get stakeholders to agree.
That is prevention advice. It is not measurement.
Nobody counts how often the goalposts actually move. Nobody tracks what the original business case said against what the program later reported. And the field is candid about its own footing: decades of research have left practitioners and scholars with a vague notion of what project success even is, settling for conflicting attributions of a phenomenon that remains elusive.
So we have a large body of work on how success should be defined, and almost nothing on what happens to those definitions once a program is under way and the original numbers are no longer reachable.
That is a strange place for a discipline to be. It is the equivalent of an extensive literature on how to set a budget, and no data at all on how often budgets get rewritten to match the spend.
Why the original number was never a fair benchmark
Based on the research, and on nearly three decades of contemporaneous observation, it is reasonable to conclude that the following is a contributing cause of the gap.
An original business case is written at the point where you know the least and must promise the most.
You know the least, because it is produced before anyone has traced how the organization actually operates. You must promise the most, because it has to clear a funding gate against competing claims on the same money.
Put those two conditions together and optimism is not a character flaw of the people writing it. It is the function of the document. The business case exists to secure a decision, and it is written by people who do not yet know what they will find.
Which means the original target and the eventual measured reality were never really commensurable. One was a persuasion instrument. The other is a measurement. Treating the gap between them as straightforward failure is a category error — and it goes some way toward explaining why organizations reach for the revision rather than the reckoning. Nobody wants to be measured against a number that was never a measurement, and should never have been treated as one.
The problem is not that targets get revised. The problem is that they get revised silently, and the original stops being referenced.
What that does to ROI
This is where it stops being a philosophical point.
If you are assessing the return on an initiative, you are comparing outcomes against a commitment. If the commitment has been restated — and the restatement is not documented, and the original is no longer retrievable — then your ROI assessment is measuring against a moving target and reporting a fixed number.
You will get an answer. It will look precise. It will mean nothing.
And this failure mode is more durable than overspending. An overrun eventually exhausts its budget and forces a reckoning. A redefined success criterion never runs out of anything. The program succeeds, permanently, against terms nobody has to defend, because the terms it was originally funded against are no longer in the room.
Which brings me to the one instance I can put a date on.
The one I can prove
In 2017 I was covering a public-sector e-procurement implementation that had been struggling for two years. I interviewed the executive who had taken over the account and asked about the adoption figures that had been the subject of my earlier reporting.
The answer was that those numbers were “not a relevant marker of success,” and that the initiative was “already a success” — having reached what was described as a critical mass of onboarding, with the remaining work characterized as optimizing usage.
I want to be careful here, because the easy reading is the wrong one. Nothing about that was evasive. It was said openly, on the record, to a journalist, for publication. She also placed responsibility for the adoption problems with the client rather than the platform — and on that point I agreed with her at the time and still do. A vendor cannot supply readiness an organization has not built, and blaming the platform for what the organization did not do is one of the more common errors in this field.
What interests me is not motive. It is that a metric can be set aside as “not relevant” without anyone lying, and the record afterwards contains no trace of the original standard. That is the mechanism, and it does not require bad faith to operate. It only requires that nobody preserved what was being replaced.
Which is why I wrote down my own test at the end of that piece: greater engagement, supplier involvement, and quantifiable taxpayer benefit — the last of those being the actual measure of success.
That post has not been edited since it was published on 3 August 2017. The publication and modification timestamps are identical, and both are held by a third party, not by me. The same applies to every post in this archive.
Nine years on, none of the three has been publicly demonstrated. The portal endures. Participation is still not universal. The system survived. That is not the same thing as the system having worked.
I do not offer that as a percentage. It is just one case example. But it is one case with the criteria written down in advance, dated by someone other than me, and checkable.
What the gap actually costs
The obvious way to price a revision gap is arithmetic: what was committed, what was delivered, subtract. That is worth doing and it is not the answer.
A business case does not merely describe a benefit. It authorizes a decision. A spend. A vendor. A sequence. A headcount plan. A set of alternatives that were considered and set aside. Every one of those was decided on the strength of a number.
If that number was never achievable, the shortfall in benefit is the smallest part of what it cost you. The real cost is everything that was decided on its strength — and none of it appears on the program’s ledger, because those decisions were booked elsewhere, to other budgets, in other years.
An overrun costs you the overrun. A revised criterion costs you the path you did not take.
And it costs you the next initiative.
People remember. A program declared successful against terms nobody recognizes teaches an organization that success is something announced rather than achieved. So the next time someone stands up to explain why this one is different, they are talking to a room that has already learned otherwise. That resistance gets called change management. Most of it is memory.
So the practical calculation is three questions rather than a formula:
What did this number authorize? Not what did it promise — what did it permit. The spend, the selection, the sequencing, the reorganization.
What else was on the table when it was authorized? The alternatives that lost. Usually documented, usually in the same paper, usually never revisited.
Would that decision have been made against the number we are reporting today?
If the answer to the third question is no, you have your figure. It is the difference between the path taken and the path that would have been taken — and it is nearly always larger than the benefit shortfall, because it compounds across every year since.
Two things follow from that, and neither is comfortable.
The first is that you cannot run this calculation without the original. Not a summary of it, not a revised version — the document as approved, in its own terms. Which is precisely what a silent revision removes. The organizations most exposed to this cost are the ones least able to compute it.
The second is that the calculation gets more expensive the longer you wait, and not linearly. Every year the original stays unexamined is another year of decisions taken on a foundation nobody has checked, each becoming the premise for the next.
And underneath those decisions, something quieter is happening.
When a system does not work the way it was promised, front-line practitioners do not stop working. They route around it. In 2005 I stood in front of a room of industrial distributors in Calgary and put the number on the screen: 79 percent of MRO purchases in Canada were going off-contract. Not because buyers were careless or undisciplined. Because getting the job done required it.
What follows is not resistance. It is concession. The organization tacitly accepts the workaround rather than examining why it became necessary — and the workaround quietly becomes the process. Nobody decides this. Nobody writes it down. There is no meeting.
A few years on, “we have always done it this way” no longer describes the system that was funded. It describes the accommodation that replaced it. And by then the original business case is not merely unexamined. It is describing something that no longer exists.
Which is the argument for doing it now rather than at the next reorganization, when the people who wrote the business case have moved on and the only surviving version of the target is the one that was met.
The fourth name
I wrote recently about a principle I had been applying for nearly three decades before discovering that other people had arrived at it in fields I know nothing about. Carl Jacobi, who inverted mathematical problems that would not yield to the forward formulas. Charlie Munger, who asked not how to succeed but how he would fail, and then stayed away from there.
There is a fourth name, and he is the one who counted.
Bent Flyvbjerg, Professor Emeritus at Oxford’s Saïd Business School and the most-cited scholar in the world in megaproject management, has spent decades assembling a database of more than sixteen thousand projects, spanning twenty fields across a hundred and thirty-six countries. From it he derived what he calls the Iron Law of Megaprojects: over budget, over time, under benefits, over and over again.
His remedy is reference class forecasting, and the logic will be familiar by now. Do not forecast from inside your own plan. Identify a class of comparable completed projects, establish the actual distribution of what happened to them, and calibrate your estimate against that distribution rather than against your own intentions.
Which is to say: do not reason forward from what you hope. Go and look at what actually happened, failures included, and reason from there.
That is one of the significant benefits of the Procurement Insights archive and the Provenance Ledger behind it — nineteen years of papers, videos, podcasts, reports and, of course, articles, dated by third parties and never edited after publication. A project database records what happened. A contemporaneous archive records what was said would happen, when, and by whom — and then what happened. The ledger exists to keep those separate, so that a recollection never quietly becomes a record.
And here are the figures that should stop anyone reading this.
Reporting on those sixteen thousand projects in How Big Things Get Done, Flyvbjerg and his co-author Dan Gardner found that 8.5 percent came in on cost and on time. Half of one percent came in on cost, on time, and on benefits.
Benefits. The dimension almost nobody hits — and the only one of the three with no external referent. A cost overrun is arithmetic; the number is the number. A schedule overrun is a calendar. A benefit shortfall is a matter of definition, and definitions can be revised.
It is worth noting how carefully Flyvbjerg and Gardner had to guard against exactly that. Their overruns are measured against the budget originally approved at the decision-to-build stage — deliberately not against figures re-baselined partway through, because business cases are re-baselined often enough that measuring against the revised number would have hidden the overrun altogether.
They had to build the defense against moved goalposts into the measurement itself, or the dataset would have told them nothing. That is independent corroboration, from sixteen thousand cases, that the practice is common enough to invalidate a study.
The half of one percent and the missing revision statistic are the same finding, approached from two directions.
What that collapse looks like at close range
I went back to something I wrote in 2008 to see whether it explained the gap between those three numbers.
It was a white paper on SAP’s Procurement for Public Sector offering, published through CATA Alliance in Ottawa. I gathered every case reference I could find — the vendor’s own published successes, the integrator case studies, and, deliberately, the failures that neither group was advertising. Roughly two dozen named organizations in total.
It is not a dataset. There is no defined population, so there is no denominator, and the two halves were assembled by opposite methods — the successes came from the vendor’s site, the failures came from going looking for failures. No percentage calculated from it would mean anything, and I am not going to offer one.
But you can see the mechanism working, and it explains the aggregate.
Of the named organizations, two claimed both cost and time. One county completed a plain-vanilla implementation with no customizations in nine months, on time and on budget. One liquor corporation reported twelve months, on time. That is roughly the same neighborhood as 8.5 percent, and at that sample size the resemblance is coincidence rather than corroboration. I note it only because it points in the same direction.
Of the named organizations, not one published what it had originally committed to deliver alongside what it delivered.
Not one. Read the success cases and the pattern is unmistakable. One city was anticipating a full return within five years — a projection, not an outcome. One council would save 72 percent of the purchase and installation cost over its first five years — future tense. One county “successfully standardized and streamlined business processes,” with no figure attached at all. One school district reported saving 50 to 75 percent of back-end processing costs, against a baseline nobody states.
These are not dishonest documents. They are simply not measurements. Every one of them reports activity, or projection, or a percentage floating free of the number it improved on.
Which is the answer to why the benefits figure collapses to half of one percent. The benefits dimension is barely reported anywhere in a form that could be checked. And a target that was never published in a form anyone could check does not need to be moved. It simply never has to be met.
You cannot move a goalpost that was never planted where anyone could see it.
The reference account
There is one case I have never forgotten, and it is the reason I stopped treating vendor case studies as evidence of anything. The white paper recorded half of it. I published the other half two years later, once I understood what I had actually been looking at.
In October 2007 I was engaged by a county government to assess a proposed business process. It was a short piece of work — days, not weeks.
In December 2007, the vendor courting that county put another county government, in a neighboring state, forward as its case reference. That is standard practice, and nobody should be criticized for it — a reference account is how enterprise software is sold. But something about it did not sit right with the people I had been working for, and they asked me to look at it.
In January 2008 — the following month — the reference county announced that after roughly two to three years and thirty-eight million dollars, its initiative had blown up, creating a political crisis.
Four weeks between being a reference account and being a catastrophe.
It would be easy to read that as deception. I do not think it was, and the honest version is more useful.
A program of that size does not collapse in four weeks. The condition existed in December. What happened in January was an announcement — the county said out loud what had been true for some time, and it made news precisely because it had not been said before. Until then, the vendor’s reference list reflected what the client was saying about itself.
Which is what a reference account is: a self-report, carrying a second party’s name. Nobody had verified it, because verification is not part of the mechanism. There is no step in that process where checking is anybody’s job.
That is a more troubling finding than deception would have been. A lie can be caught. This cannot, because nothing went wrong — every party behaved exactly as the process expects, and the output was still worthless.
When I wrote about it again in 2010, I put it in one line, and I have not improved on it since: today’s case reference could be tomorrow’s embarrassment.
That is what a success set composed of self-reported cases with no verifiable benefits data is actually worth. And it is worth remembering when your own program is being assessed on the strength of somebody else’s reference list — or when yours is about to become somebody else’s.
Four operations, one principle
I want to be precise about what is shared and what is not, because a loose version of this claim is worth less than an accurate one.
These are four different operations, and the differences are not cosmetic.
Jacobi restates. He takes a problem that resists a frontal assault and flips it into its inverse form, where the path becomes visible. The work is mathematical and it is done on paper.
Munger avoids. He imagines a failure that has not occurred, enumerates the ways it could arrive, and steers clear of them. The failure is hypothetical by design; the point is never to meet it.
Flyvbjerg counts. He takes failures that have already finished, aggregates them across an entire class, and uses the distribution to correct a forecast before a new project starts. The failures are real but they are other people’s, and they are closed.
I trace. The failure is not hypothetical, not aggregated, and not finished. It is happening right now, inside an organization, with people in it who have positions attached to how it gets explained. I start at the breakage and follow the mechanism backward until it terminates in a cause — which is routinely somewhere nobody was looking, and frequently outside the department that owns the problem.
That last difference is the one that matters in practice, and it is not a claim of rank. It is a claim about conditions. Three of these operations can be performed with the evidence sitting still. Mine has to survive contact with an operating reality that pushes back while you are working on it — where the people who chose the original path are in the room, and where the cause you are tracing toward may be sitting across the table.
Underneath the four operations, one idea: some problems cannot be solved forward.
I did not learn it from any of them. I developed and proved my method in 1998, with funding from the Government of Canada’s Scientific Research & Experimental Development (SR&ED) Program, having never encountered Jacobi’s maxim and knowing nothing of Munger’s work. You cannot operationalize what you have not read.
Which is the most persuasive thing about it. A preference held by four people who share a reading list is a school. A principle that surfaces separately in pure mathematics, in investing, in the statistics of megaprojects, and in enterprise failure diagnosis — in four operations that do not resemble each other — is not a preference. It is closer to a law.
What the backward approach does to ROI assessment
This is the practical payoff, and it is the reason any of the above matters to someone with a live initiative.
An equation-based approach starts from the goal and builds toward it. Its defining property is that it commits you, structurally, to being right about the path you chose. Every subsequent fact sorts into helps me get there or obstacle. You are no longer investigating; you are defending a destination. And when the destination becomes unreachable, the cheapest available move is not to abandon it — it is to redescribe it.
An agent-based approach starts where the thing is failing and traces backward. Because you began at the breakage rather than the goal, you have nothing to defend. You can follow the evidence out of your own department, past your own decisions, to a cause nobody suspected.
Applied to ROI assessment, that produces three things the forward method cannot:
A retrievable original. The trace forces you to establish what was actually committed to, because that is where the trace terminates. You cannot quietly lose a document you are working backward toward.
A measurement against the operating reality rather than the pitch. The trace tells you which aspects of the operating reality were validated and which were only assumed. Return calculated against validated conditions means something. Return calculated against assumed conditions is arithmetic performed on a wish.
A documented delta. This is the deliverable that matters, and almost nobody produces it: what was assumed, what the trace found, what changed, and why. Not a replacement business case. The original and the revision, side by side, with the evidence between them.
A real and documented initiative, mapped two ways. On the left, the process as designed — every arrow pointing at the target, containing only what someone thought to build in. On the right, the trace: it leaves the mapped process at the second step, loops on itself, and terminates in a cause nobody had thought to include. First published in The Map You Design and the Map You Trace, 29 July 2026.
The test
Because targets should sometimes change. An initiative that discovers its founding assumptions were wrong and carries on toward the original number is not disciplined; it is stubborn. Revision can be the correct output of good diagnostic work.
So how do you tell the difference between a legitimate revision and a moved goalpost?
Three conditions, all required:
The third is the one that does the work. A revision can be announced openly, on the record, for publication — and still be a moved goalpost, if the original is not preserved and the change is asserted rather than derived from evidence. Openness alone does not qualify it. In the 2017 interview above, the new criterion was stated publicly and without evasion. The original was simply set aside as not relevant. Nothing was hidden. The goalpost moved anyway.
Which is why the missing statistic may be missing for a reason. You cannot count what nobody is required to preserve.
If your organization is running an initiative right now, there is one question worth asking this week, and it does not require an assessment, a consultant, or a methodology:
Can someone produce the original business case — and do the measures we currently report against still map to it?
If the answer is yes, you are in better shape than most.
If the answer is difficult to find out, you have just learned something more valuable than any ROI figure your current reporting would have given you.
-30-
Truth Is Believing. Accuracy Is Knowing. Outcome Is Proof.™
Share this:
Like this:
Related