The Agentic Organization Isn’t New. It’s a Bigger Trainyard — and the Real Risk Is an Empty Tower.

Posted on July 24, 2026

0


Truth Is Believing. Accuracy Is Knowing. Outcome Is Proof.™

McKinsey and others are naming the agentic organization — an enterprise where AI agents take on real decisions. The naming is useful. The structure it describes, I ran a version of in 1998. What has changed is the size of the yard. What has not changed is who has to be in the tower.


The agentic organization is having its moment, and the moment is earned. When McKinsey and others describe enterprises where AI agents carry real operational decisions — pricing, sourcing, scheduling, replenishment — inside defined authority, they are describing something that is genuinely arriving. I am not here to argue it away. I am here to say the shape of it is not new, and the part of it that decides success or failure is the part the excitement tends to skip.

I know the shape because I built one in 1998.

The trainyard

The engagement was at Canada’s Department of National Defence — early work funded in part through the Government of Canada’s Scientific Research and Experimental Development program — and the problem was the timely acquisition of MRO parts against a national IT infrastructure. The architecture I built to solve it, RAM 1998, worked like a trainyard.

On the tracks were the buyers. Each one took the system’s supplier ranking — scored on historical performance and real-time parameters — and applied their own weighting to it. Prioritize delivery, and the suppliers re-ranked on delivery history; the lowest bid did not automatically win. Prioritize price, and they re-ranked again. That was human judgment in the loop, exercised on each track, by the person closest to the work.

Above the tracks was a central manager who could see the whole yard in real time — every buyer’s activity, at any moment — with the authority to step in. That layer caught two things nothing else would: the model drifting away from reality, and the buyer quietly gaming the model toward a worse outcome. That was human judgment over the loop.

Two layers, deliberately distinct: discretion on the track, governance in the tower. On-time delivery moved from 51% to 97.3% in three months — and the record building itself as the trains ran was the point, not a footnote. I will come back to that.

What changes in 2026 — and what does not

Put the agentic organization next to the trainyard and the mapping is exact.

The trains on the tracks become agents. They can make the price-and-delivery calls a buyer used to make — faster, continuously, at a scale no room of people could match. That part is real, and it is worth having.

The human does not leave the yard. The human moves up. In 1998 I was on a track. In the agentic organization, the human belongs in the tower — governing a yard that now contains agents rather than only people. The role does not disappear as capability rises. It elevates, because governing capable agents demands more judgment than doing the task yourself, not less.

That elevation is the whole claim, and it is the core of Invariant Physics™: the technology under your hands changes endlessly — ERP, the internet, cloud, mobile, and now agents on the tracks — while the requirement that a human govern the operating reality against a documented record does not move. This is not a prediction, and I want to be precise about that, because prediction is the lazy version of the claim. It is a falsifiable proposition. Show me an agentic rollout that succeeded with the tower empty — no governed boundaries, no human owning the outcome — and it counts against the principle. Across the record, back through every wave to that 1998 yard, that case has not appeared.

The stress point the metaphor exposes

Here is where the trainyard earns its keep, because it surfaces the exact thing the agentic conversation gets wrong.

In 1998, “the manager monitors the yard in real time” was a promise the design could keep. A person can watch a room of buyers moving at human speed. But a yard that is exponentially larger, with trains running at machine speed, cannot be watched that way — by anyone. Pretending it can is abdication wearing a hi-vis vest: it looks like governance while being physically impossible.

So the tower’s job changes shape, not just scale. You govern by boundary and exception. You own the rules each track runs under. You own the exception-flags that surface only the trains that need a human. And you own the outcome. What you cannot do is personally watch every train — and the honest version of this says so out loud.

That has a consequence people would rather not face. If the yard is too large to watch, something has to watch the tracks for you — and increasingly that watcher is itself AI. The moment that is true, the human’s job becomes seeing into and auditing the watcher. A manager who cannot inspect his own monitoring layer is not governing. He is trusting. That is the difference between transparency and theater, and it is the entire reason a verification discipline — multi-model, adversarial, and traceable to a primary record, the work I now run as RAM 2025™ — exists. Not more surveillance. Auditable governance.

Why this isn’t a dashboard

By now someone is thinking: this is a dashboard, and there are plenty of those. It isn’t — and the difference is not cosmetic. It is the whole point.

Two things separate a trainyard from a dashboard.

The first is what’s on the tracks. A dashboard is a display bounded by the enterprise — a rear-view pane showing you the state of what you already own. The trainyard is not something you look at; it is the governed system itself, and its agents are both human and machine, running past the enterprise wall to the suppliers, the couriers, the customs broker. In 1998, the agents whose behavior determined the outcome were never all inside the building — the parts came from US suppliers, across a border, through customs. That extended network of agents is what I have long called the Metaprise™. A dashboard shows you your yard. The trainyard governs a yard that was never only yours.

The second difference is the one that actually matters, and it turns on a distinction almost no one makes. Traditional data cleaning removes malformed data. It cannot remove valid records of invalid behavior.

The lesson is not “greenfield wins” — no enterprise gets to un-have its lake. The lesson is that point-of-capture validation is the discipline separating a learning system from a stored one, and it is a tower function you can impose now. The tower doesn’t only watch the trains. It decides which data the agents are allowed to learn from — and that is governance, not display. No dashboard does that.

I will go deeper on the difference between the trainyard and the dashboard model in an upcoming post.

Why the bigger yard is survivable now

There is one asymmetry that makes an exponentially larger yard governable rather than terrifying, and it is the opposite of what you would guess.

In 1998, the yard started empty. There was no prior record of comparable conditions when the system switched on. The trainyard built its substrate as it ran — the climb from 51% to 97.3% was the record accumulating in real time, one order at a time.

In 2026, the human walking into the tower does not start from zero. For the first time, the manager arrives with decades of documented pattern already on the desk — against which today’s machine-speed decisions can be checked. The yard is bigger. The record is deeper. And because that record was built the way the 1998 yard was built — validated at the point of capture, not scraped from whatever the transactions happened to log — it is a substrate the agents can safely learn from rather than a lake that would teach them the wrong thing. That is precisely what makes the tower a viable place to stand: the substrate that took 1998 three months to begin building is already there before the first agent moves.

The danger, stated plainly

The mistake in the agentic-organization conversation is the extreme — the quiet assumption that sufficiently capable agents let the organization step back and let the yard run itself. The danger was never capable trains. The danger is an empty tower.

Transparency and governance are the tower. Remove the human from the boundary that decides which decisions agents get to make and which get escalated — remove the accountability for the outcome — and you have not gained autonomy. You have lost the yard.

There is a quieter version of the same abdication, and it is already underway. Point capable agents at a data lake full of valid records of invalid behavior, with no tower validating what they learn from, and you don’t automate judgment — you automate the dysfunction, at machine speed. The failure pattern doesn’t just repeat; it compounds. An empty tower and an unvalidated substrate are the same mistake wearing two hats: in both, no human owns the boundary between what was recorded and what is real.

I am not the only one pointing at this. Under McKinsey’s own post, Marian Nagy — who leads enterprise architecture in banking — put it plainly: “the primary AI risk isn’t technological failure; it is the scaling of organisational weaknesses.” AI, he noted, will not hide the data debt, shadow applications, legacy systems, and siloed data an organization already carries — scaled across the enterprise, it amplifies them, unless the company continuously assesses its true readiness across process, data, and infrastructure first. That is the trainyard’s conclusion reached from the other direction: the danger is not the capability, but what it is pointed at — and whether anyone is governing the aim.

The agentic organization was always going to arrive. The naming is overdue and the capability is real. The only open question is the one the excitement keeps skipping:

When it arrives, is anyone in the tower — and can they clearly and accurately see the tracks?


This analysis draws on the Procurement Insights archive — an independent record I have published openly since 2007, consolidating documented client work, lectures, and writing reaching back to 1998, and carrying no vendor sponsorships across the past decade. Every claim is held to the Provenance Ledger™, a verify-before-publish discipline that traces each assertion to a primary source and reconciles the record forward rather than editing it in place. Invariant Physics™ is the constant it keeps testing; Implementation Physics™ is its per-engagement application, run through Phase 0™. Getting it right, rather than being right.

Related: The Biggest Mistake People Are Making with AI · 1,800 Hours With AI, and the One Thing That Never Changed

-30-

Posted in: Commentary