There Is No Best AI Model. There Are Only Different Models — And You Have To Choose Your Team Wisely.

Posted on September 13, 2026

0


Dr. Patrick Giwa posted something this week that is worth reading carefully, because he is right about almost all of it.

His point: every few weeks there is a new model update, a new benchmark, a new leaderboard, a new post declaring the obvious winner. Three months ago the advice was to move to one provider. This month it is to move back. Meanwhile, benchmarks matter less than how a tool performs in actual work, switching after every update costs more than anyone admits, most professionals use a fraction of what they already have, and — his best line — a better model will not fix a poor workflow.

I agree with the direction. I would put a different question in front of it.

The question underneath

“Which model is best?” is the wrong question, and not because the rankings keep changing. It is the wrong question because it assumes a single answer exists.

I have been in high technology since 1983, and I have worked with advanced self-learning algorithms in a nascent AI platform since 1998 — a Department of National Defence engagement, funded in part through Canada’s SR&ED program, where on-time delivery moved from 51% to 97.3% in three months. There was no model to pick in 1998. There was a system, an organization, and the gap between what the system assumed and how the work actually happened.

Today there are models to pick, and the picking has become the conversation. That is the part I would push back on.

There is no such thing as a best model. There are different models — with different strengths, weaknesses, behaviors and blind spots. Much like people, the question is not which individual is best. It is which team you assemble, what role each member plays, and how you challenge what comes back.

Leaderboards rank models. Real work requires choosing a team.

Procurement has seen this before

If that sounds familiar to anyone in procurement, it should.

We have spent two decades watching organizations select technology from analyst rankings. In January 2011 I wrote about a firm being named a leader in supply chain planning and asked what the Veterans Health Administration would have had to say about it. The ranking was not false. It was simply a measure of the vendor, not a measure of how the thing would behave inside a particular organization.

That is exactly what a model leaderboard is. It tells you something real about the model. It tells you nothing about the fit between that model and your work — which is why the Hansen Fit Score™ measures capability against demonstrated outcome rather than ranking suppliers against one another. A 7.0 on capability and a 1.5 on outcome is not a contradiction. It is the whole story.

Swap “vendor” for “model” and the failure mode is identical. Same mistake, new category.

What I found when I stopped picking

I have put more than 2,000 hours into structured work with these systems — single-model and multimodel, always with a human orchestrator in the loop. I wrote up the findings in July. Four of them matter here.

Confidence is not correctness. These systems produce answers that are articulate, fast and assured, and a meaningful share of them are wrong in ways the confidence conceals. Not obviously wrong. Plausibly wrong.

The most agreeable voice is the one to watch. In a multimodel setting they do not behave identically. Some are cautious, some combative, some agreeable. The agreeable one feels best to work with and is the most dangerous to trust. The comfort is the tell.

Convergence is not consensus. When independent systems reason separately and land on the same load-bearing point, that is signal. When they round toward the same inoffensive middle, that is averaging — and averaging sands the edge off the arguments that were worth making.

Friction is the point. A panel is not for more answers. It is for productive disagreement. The human brings the signal; the systems provide the friction. Neither substitutes for the other.

Notice that not one of those findings is about which model is best. Every one of them is about roles, temperament and the human sitting in the middle of it.

That is team construction. It is also why, in Thinking With the Machine, the models are identified by number rather than brand. The roles they play matter. The logos do not.

Where Patrick’s advice becomes load-bearing

His closing sequence is sound: know the work you want to improve, pick a strong tool, build repeatable ways of using it, switch when there is a clear reason.

I would add one question ahead of the first step.

How do you know the work operates the way you believe it does?

Because repeatability is only an advantage after the operating model has earned the right to be stabilized. Make a workflow repeatable too early and consistency stops being a strength — it starts preserving the error. You can become model-independent, build a highly repeatable workflow, execute it consistently, and get very efficient at doing the wrong thing.

In 1998 nobody needed a better tool. Everyone involved was behaving coherently inside their own frame, and every measure reconciled. The breakthrough came from a question outside the accepted frame: what time of day do the orders come in? Only then did the operating relationships become visible.

Good workflows survive model changes. Good workflows built on untested assumptions survive model changes too. That is the problem.

A 30-minute lab

I am running a free 30-minute session on this, and it is a working lab rather than a presentation.

You will be handed a real, dated exchange — a public post about model selection, a panel model’s advice on how to answer it, and what was actually posted — and asked to make the same calls I made, before you see what I did.

Somewhere in the middle, one of the models takes a sharp comment of mine and makes it smoother. Your job is to notice — and you will leave with the question that catches it: is this better, or is it just smoother?

What you leave with: A better way to select, work with, and progressively engage with AI models to optimize output accuracy and deliver quality insights.

Put another way: this session and the labs teach the same thing — how to reason toward a successful outcome by leveraging AI, rather than asking AI for the answer.

And the method has a track record. In 1998, RAM selected the right supplier 97.3% of the time — measured the only way that counts, on outcome. On-time delivery moved from 51% to 97.3% in three months and held there for seven years. Across more than 2,000 hours of structured multimodel work, ARA™ RAM 2025™ runs at 91%. The technology is unrecognizable. The method has not changed.

Friday 25 September, 9:30 AM Eastern. Free, and everyone who attends gets a copy of Thinking With the Machine.

Register here: https://www.linkedin.com/events/7504567838724997120/

A second run follows on Wednesday 30 September at 1:00 PM Eastern, for anyone the Friday slot does not suit.

It is time to start working with technology rather than for technology.

-30-

Truth Is Believing. Accuracy Is Knowing. Outcome Is Proof.™

Jon W. Hansen, FCIPS — Procurement Insights | Hansen Models™

Posted in: Commentary