Everyone Is Sharing the AI Toolkit. No One Is Asking What’s Underneath It.

Posted on June 1, 2026

0


There is an infographic circulating in procurement circles right now, and it is genuinely useful. It maps how procurement teams can use Microsoft Copilot: Excel to analyze RFQ responses and flag the lowest-cost supplier, Word to summarize contracts and draft supplier emails, PowerPoint to turn insights into executive slides, a chatbot to generate negotiation strategies. Each cell comes with a ready-made prompt. It is clean, practical, and people are right to share it.

It is also a near-perfect picture of the thing that has failed in every technology era before this one — and the reason it will fail again is not in any of the cells. It is in what the infographic does not show.

Agents at the frontline are not an agent-based model

Look closely at what that toolkit actually describes. It is a set of individual people, at their own desks, handing discrete tasks to an AI agent and acting on what comes back. Summarize this contract. Rank these suppliers by cost. Draft these negotiation points. Each one is a frontline worker invoking an agent, in isolation, with no shared architecture beneath them.

That is technology-led agentic activity. It is not an agent-based model, and the difference is the whole game.

An agent-based model is built from the real-world operating attributes of the actual stakeholders — buyers, suppliers, couriers, the people who live inside the process — and it captures how their behaviors and incentives interact to produce outcomes. In practice, that means modeling things like supplier reliability patterns, inbound-logistics failure points, plant-level constraints, and the incentive structures that actually drive behavior — before any agent picks a supplier. The agents in that model are grounded in operating reality. The agents in the Copilot toolkit are grounded in nothing but the prompt and whatever the open document happens to contain. One is an architecture. The other is a thousand disconnected queries.

The toolkit teaches the second and calls it AI adoption. And because it produces fluent, confident, useful-looking output at every desk, no one notices that the foundation the output should be standing on was never built.

Why “highlight the lowest-cost supplier” is the tell

Take the most innocent-looking prompt on the entire infographic: analyze this RFQ sheet and highlight the lowest-cost supplier.

It will work. It will return an answer, instantly, with confidence. And in a great many real procurement situations it will be wrong — not because the model failed, but because the lowest quoted cost is rarely the lowest landed cost. The variables that actually determine the outcome — total cost of ownership, delivery reliability, quality, the conditions upstream and downstream of the purchase order — are not on the sheet the agent is reading. They live in the operating reality the agent has no access to.

A buyer who acts on that output has not made a better decision faster. They have made a confident decision on incomplete ground, and they have done it invisibly, with nothing in the organization positioned to catch it. Multiply that by every desk running the toolkit, and you do not have an efficiency gain. You have unvalidated agentic output flowing into real decisions at scale, with no one watching.

This is the absorption gap, materializing one siloed event at a time. What you get is a collective of agents, not an aligned enterprise — and the two are not the same thing.

The intervention happens before the AI arrives, not after

This is where my position parts company with most of the current conversation about AI. The debate today is almost entirely about what to build after deployment — governance layers, orchestration, decision rights, oversight to catch the AI when it drifts. All of that is real, and all of it is downstream. It treats the missing piece as something you bolt on once the agents are already running.

The intervention that actually determines the outcome happens earlier — before the AI exists. It is the work of understanding the operating reality first, and building the model on it, so the agent is grounded from the start rather than corrected after the fact. The data an AI agent produces is only as good as that prior understanding. Not “a person checks the answer” — but the model built on real operating conditions, so the output reflects how the thing actually works rather than how it appears on a spreadsheet. Post-deployment governance inspects output. Pre-deployment understanding shapes whether the output was ever worth producing. They are not the same intervention, and only one of them changes the result.

It is worth pausing on one persistent belief here: that “mapping the process” already does this work. The standard move is to interview stakeholders in each department and chart an equation-based pipeline flow. On the surface, that looks like the right approach to automation. But it rarely captures the reality of an organization’s true operating environment — least of all the internal and external stakeholder incentives that actually determine the outcome. I wrote about exactly this gap in a case the process map would never have surfaced: Tell me how Agentic AI would have known to ask, “What time of day do orders come in?” The answer the question forces is the whole point — no flow diagram contained it, because it lived in the incentives, not the flow.

That kind of intervention is hard. It is precisely the work most organizations skip, because it is slower and less satisfying than typing a prompt and getting a polished paragraph back. The Copilot toolkit does not just omit that intervention — it makes omitting it feel like progress. It hands every frontline worker the ability to generate authoritative-looking output without ever touching the conditions that would make the output trustworthy.

This is not a criticism of the people using it. It is a criticism of an approach that gives them powerful agents and no architecture to ground them in — and then measures success by how much the agents are used rather than whether the decisions held.

The pattern does not change

I documented this position in 2004 and again in 2007: real control of a spend environment comes from a decentralized architecture built on the frontline operating attributes of every stakeholder, not from a central tool issuing answers. The method that proved it was developed in 1998 with funding from the Government of Canada’s Scientific Research and Experimental Development program (SR&ED), and the results held for seven years — 23% year-over-year savings, the collective buyer count reduced from twenty-three to three — because the model was built on operating reality first, and the technology was introduced last.

Unfortunately, every era since has offered a new version of the same process mapping shortcut. ERP, SaaS, digital transformation, and now agentic AI at every desk: each promises that the tool will deliver the outcome, and each produces the same gap when the tool is deployed without the architecture beneath it. The acronyms change. The missing layer does not.

The Copilot toolkit is not the problem. It is a capable set of tools, and on top of a real agent-based model it becomes an accelerant rather than a risk. The problem is what is being shared with it today — which is nothing. No model of the operating reality. No agent-based architecture. No account of the conditions that decide whether “the lowest-cost supplier” is actually the right answer.

Before you hand your team the toolkit, it is worth asking the only question that has ever determined the outcome: what is underneath it? If the answer is a prompt and a hope that the document contains everything that matters, you do not have an AI advantage. You have shadow work with better grammar — and a failure rate you will recognize the moment the numbers come due, only sooner now that AI is the one producing it.


Based on observations documented continuously in the Procurement Insights archive, including agent-based modeling work first published in 2004 and the SR&ED-funded engagement that proved it in production. Archive since 2007; practice and proof lineage to 1998. Zero vendor sponsorships.

-30-

Posted in: Commentary