When Is the Right Time to Disturb Success?

Posted on August 26, 2026

0


Everyone knows what to do when the numbers are bad. Almost nobody has a discipline for what to do when the numbers are good and the operation is still wrong.

Jon W. Hansen, FCIPS | Procurement Insights | August 2026


THE SHORT VERSION FOR BUSY EXECUTIVES

There are two moments in an operation when something can be found.

The first is when something is visibly failing. That moment is uncomfortable and it is also generous: the failure is evidence, it demands an explanation, and nobody has to be persuaded that looking is worthwhile.

The second moment is when the numbers are clean. No signal fires. Every party is performing. The contract is safe. And the operation can still be wrong — because a measure can be accurate and be measuring the wrong thing.

The second moment is where the expensive findings live, and it is the one nobody is organized to work. There is no budget line for come back when things are going well and tell us what we are still getting wrong.

I have one case where both moments occurred in the same engagement, three months apart. The second was worth more than the first, and nothing in the system asked for it.

A DEEPER DIVE


The first moment: a failure that demanded an explanation

In 1998 a national defence maintenance operation was delivering next-day parts on roughly half its requirements against a 90 percent contractual requirement. The contract was at risk.

The declared problem was procurement, and the declared fix was to automate it.

I asked a question that was on no map: what time of day do orders come in?

Four in the afternoon. Asking why took me out of procurement entirely, into a service department where technicians were holding parts orders to the end of the day because they were rated on the number of calls they responded to. Every one of them was doing exactly what they were measured on, all day, competently.

That behaviour set everything downstream in motion, and I have traced the full route elsewhere.

Within three months delivery ran above 97 percent — before any new technology entered the building.

None of that was hard to justify. The operation was failing. A failure is the one thing in a system guaranteed not to have been designed, which makes it the only reliable place to start. Nobody had to be convinced that looking was worth the effort.


The second moment: nothing was wrong at all

Some months later, with the system in production and delivery holding above 97 percent, I went looking again.

There was no reason to. The contract was secure. Every measure was green. The client was satisfied and had every right to be.

What I found was this. A part that arrived on time counted as success. A part that arrived on time and was the wrong part also counted as success.

On-time delivery of the wrong item had been scoring as a win for the entire life of the engagement, and nothing in the instrument could distinguish the two.

So we added an attribute the original design had never specified: capturing defective and wrong-part shipments. Once correctness and quality wrote back to the same record as delivery, a supplier who shipped fast and shipped wrong lost standing on the very instrument that had been rewarding them.

That finding could not have surfaced from the failure. The failure had already been explained and corrected. The delivery number was clean. There was no signal, no anomaly, no complaint — nothing fired.

It came from a decision to look where nothing was flashing.


Why this is the harder discipline

The first moment has a mechanism behind it. A delivery rate stuck at half against a 90 percent requirement produces itself, whether or not anyone remembers a methodology. It cannot be argued out of existence, and it does not depend on anyone’s attention.

The second moment has no mechanism whatsoever.

Nothing tells you it is time. No threshold is breached. No exception is logged. The organization is performing, the reports are accurate, and every incentive in the system says leave it alone.

Which means that at the exact moment an operation is most exposed to a wrong measure — when the measure is producing good numbers and therefore nobody is examining it — scrutiny is at its lowest.

Engagements close when the metrics improve. Reviews move from weekly to quarterly. Dashboards go green and stop being read. The clean number is treated as an answer when it is only a reading.


And it is about to matter considerably more

An agent deployed against a measure inherits that measure’s blind spot and executes against it faster than a person could.

If on-time delivery is the measure, an autonomous system will optimize on-time delivery — including for the wrong part — with more consistency and less hesitation than any human buyer, and every indicator on its dashboard will improve while it does.

A system reasoning from a wrong measure does not produce a visible error. It produces confident, well-executed, wrong outcomes at scale.

The reported figures on AI implementation now bear this out — though the most-quoted one deserves more care than it usually gets.

RAND’s study of AI project failure opens by relaying other people’s estimates: more than 80 percent of AI projects fail, roughly twice the rate of IT projects without AI. RAND is characterising figures it did not measure, and its own method was interviews with sixty-five data scientists and engineers. Read it as the large majority fail, not as a percentage — and treat any version carrying a decimal place as a sign that whoever wrote it did not read the source.

What RAND did measure is the part that matters here. From its own interviews, the root causes are overwhelmingly leadership and organizational rather than technical: misaligned purpose, weak data foundations, integration as an afterthought, fading sponsorship. Almost none concern the algorithm.

MIT’s Project NANDA puts around 95 percent of organizations at no measurable return from generative AI pilots.

Notice what those studies measure: whether the outcome moved. Earlier technology eras counted a project successful when it went live. The measure got stricter, and the number got worse.


So when is the right time?

Not when something breaks. That moment takes care of itself.

The right time is when you have just succeeded — when the number came good, the pressure came off, and the temptation to stop looking is strongest. That is the moment the model underneath the number has been least examined and has had the most opportunity to drift from the operation it describes.

Three questions worth asking while things are going well:

What does this measure count as success that should not be? On-time delivery of the wrong part counted as a win for the whole life of the engagement to that point.

Which party’s behaviour determines this outcome, and is that party measured on it? In 1998 the answer was a service technician who had never been asked about parts delivery.

When was this measure last examined — as opposed to reported? Reporting a number and testing whether it still means what it meant are different activities, and only one of them appears on a schedule.


The honest version of the argument

I am not claiming every clean number conceals a problem. Most do not. An operation whose declared process matches how the work actually runs is fine, and examining it is wasted effort.

The claim is narrower: you cannot know which case you are in without looking, and nothing in the system will tell you it is time.

That is what makes it a discipline rather than a procedure. A procedure fires on a trigger. This one has no trigger — which is precisely why it has to be deliberate, scheduled, and undertaken by someone whose standing does not depend on the number staying good.

Nobody disturbs their own success unprompted. There is no reason to, and every reason not to.


Jon W. Hansen, FCIPS — Procurement Insights | Hansen Models™ | Independent. Unsponsored. Archive-based.

Truth Is Believing. Accuracy Is Knowing. Outcome Is Proof.™

-30-

Posted in: Commentary