Score the Decision, Not the Technology

Posted on September 8, 2026

0


Yesterday’s post argued that the failure rate has not moved in twenty-five years because the sequence has not moved. Capability first, conditions later, assessment after the commitment.

A reasonable objection follows. Everyone believes they assess before they commit. Nobody funds a program intending to skip the readiness question. So if the sequence is the problem, why does it survive people who are actively trying to avoid it?

I have spent the last two days inside a small version of that question, and the answer is not what I expected.

What the lab actually produced

The exercise began with an anomaly in this blog’s own traffic. One country’s numbers departed sharply from a stable eight-year pattern, and I wanted to know why.

The work ran through a structured multimodel environment against two analytics platforms, neither of which I control and neither of which can be opened. The record runs to roughly eighty pages, and it is a transcript rather than a report — every reversal, every correction, and both of the substantive errors are in it, dated, in the order they occurred.

I am not publishing the figures. They are not the finding, and after two days of work I would not put them in front of anyone making a decision. That is the finding.

Here is what happened.

Every model produced analysis that was fluent, internally consistent, and appropriately hedged. Two of those analyses were wrong in ways that changed their conclusions. One misread a single figure and built a narrative of collapse on it, along with a plausible hypothesis about why. Another — mine to correct, and I did not catch it either — applied three different interpretive standards to the same class of evidence, and reached a conclusion about audience that depended on a step it had not earned.

Both errors were caught. Neither was caught by the model that made it, and neither was caught by consensus. Adding a sealed second panel whose only task was to attack the first found the one the main panel had missed.

And across eight models and two days, not one of them asked the question that governed everything else: how does this instrument produce these numbers, and is it good enough for the decision I am about to make?

I asked it at the end. The answer was that I do not know how either platform attributes what it attributes, I cannot verify it, and the confidence I would place in the output for any decision involving money is about three out of ten.

At that point the correct move was to stop. Not to gather more data. Not to run another panel. To stop, because no amount of additional analysis improves an instrument you cannot inspect.

Why that is not a story about analytics

An organization evaluating an AI initiative in 2026 is standing in the same position, and most of them have not noticed.

The performance data comes from a platform they do not control. The evaluation frameworks come from firms with a commercial interest in the category. The models advising them will produce confident, coherent, well-structured analysis whether or not the underlying evidence merits it — that is what they do, and it is not a defect. The output looks the same either way.

What almost never happens is somebody asking what the analysis rests on before the analysis is believed. That is why clean data on its own is insufficient. It is not what your data contains. It is what your data cannot, and often does not, see.

In short, you do not just need clean data. You need the right data — and the right data is usually sitting at an unowned boundary, or several.

That is the same failure as the 1998 case. Every measure on the board was accurate. The determining condition sat outside the frame, and no amount of accuracy inside the frame was ever going to reach it.

So: score the decision

The question I keep being asked is what score I would give an AI initiative in 2026.

It is the wrong unit. The technology does not have a score. The decision does, and the score depends entirely on what is known at the moment the commitment is made.

The decision in front of youProceed?
Buy or scale AI because most organizations already have2
Pilot the popular use case, prove value later3
Conditions unknown, vendor already chosen3
Three questions answered, bounded first slice, review date set6–7
Measured gain achieved before the platform decision7–8 on that slice — not on the enterprise program

The threes are not my opinion. They are the published scoreboard for this wave, inverted: roughly a fifth of organizations have scaled AI across business units, around two-fifths report measurable bottom-line impact, and a single-digit percentage report significant value. Those figures come from three different firms counting three different things, so they are not a series — but nothing in that range supports a confident commitment.

The six to seven is my judgment, and it should be read as judgment. There is one supporting data point worth knowing: in the same survey that produced the scaling figure, organizations that track return continuously, treat AI as a portfolio, and actually discontinue initiatives that are not working reported positive returns on the large majority of them. Different measurement, so it does not net cleanly against the first — but the direction is unmistakable, and the differentiating variable is not the technology, the vendor, or the use case. It is whether the organization can measure what it is doing and stop something that is not working.

What moves the score

Three questions, answered before the commercial commitment rather than during implementation.

What determines the outcome here, and is it currently being measured? In 1998 the answer was what time of day the technicians submitted their orders, and it was on nobody’s scorecard.

Where does the process cross a boundary that nobody owns? Handoffs are where value dies unobserved, because no single function is accountable for the crossing.

What would have to be true for this to work, independent of which product is selected? If that question cannot be answered without a brand name in the sentence, the assessment has not happened yet.

Put the two artifacts side by side and the difference is not subtle.

[GRAPHIC — The Map You Design vs. The Map You Trace] Source: The Map You Design and the Map You Trace, Procurement Insights, 29 July 2026.

The map on the left is the process as it was described — five steps, every arrow pointing at the target, ending in a green check. It contains only what somebody thought to design into it, which is why it can be entirely accurate and still miss the thing that decides the outcome.

The map on the right is the same operation, traced backward from the failure. It leaves procurement four times — into service, back into procurement, out again at the border, then into finance — and each of those crossings is a boundary nobody owned. Note also where four o’clock sits: in the middle of the chain, not at the end of it. It was a correct answer. It was not the last one.

None of that requires a large program. It requires somebody with standing to ask, and an answer that is allowed to be no.

The part that is hardest to hear

A three dressed up as an eight is what frequently gets funded.

Not through anyone’s bad faith. A three is unfundable — no career advances by proposing restraint and no budget line pays for it — so the case that reaches the committee is the one that reads as an eight. Everyone in the room is behaving reasonably, and the sequence survives another generation.

Refusing that dress is not analysis. No model does it, and the two days behind this post are the evidence. It is a judgment about whether the evidence clears the bar for the decision being contemplated, and that judgment sits with a person.

You would not invest against a three. Do not capitalize a program on a slide that cannot name the determining condition, the unowned boundary, or what would have to be true independent of product.


Truth Is Believing. Accuracy Is Knowing. Outcome Is Proof.™

-30-

Posted in: Commentary