A cone does not know whether it is in the right place. It only proves somebody put it there.
Truth Is Believing. Accuracy Is Knowing. Outcome Is Proof.™
Executive Quick Read: Guardrails close on the output; RAM closed on the operation.
Adaptive guardrails exist and are deployed. They adapt to the wrong thing — to attack probes written by the vendor’s own red team, to the model rather than to themselves, or to whatever a human expert last wrote into an evaluator. None of them closes on what happened in the operation.
The research literature says these controls are typically specified before deployment and remain largely fixed during use, and that the response to a real-world failure is to collect examples and retrain. Read as a sequence, that means the guardrail is always one incident behind. Survivable at quarterly review speed. Different when an agent is acting continuously on the unfixed condition.
A guardrail that closes on the output can pass every response in a programme that is failing. In 1998 the determining condition in a defence maintenance operation was a technician incentive in another department — nothing in any output would have shown it.
So the question for anyone selling AI governance is not whether the guardrail adapts. It is what it adapts to.
In April 2025 I looked at how five ProcureTech providers described their agentic AI, and made a general notation:
None of them provides a tangible description or explanation of the learning loopback process for their self-learning algorithms — which is the most crucial component of an effective Human-Agentic AI dynamic.
Eighteen months later, the same test applies to a different subject.
Nobody has described how a guardrail learns.
What the record actually says
This is not a hunch. The field says it about itself.
In May 2026, researchers at Google Cloud AI Research and Seoul National University published LiSA: Lifelong Safety Adaptation via Conservative Policy Induction. Surveying the guardrail field, they find these controls rely on a static, general-purpose definition of harm specified before deployment, and remain largely fixed during use. And they name why that is hard to fix:
The hardest failures are often contextual: whether an action is acceptable depends on local privacy norms, organizational policies, and user expectations that resist pre-deployment specification.
The conditions that decide whether a control should fire cannot be listed in advance. That is not my claim about the industry. It is the industry’s claim about itself.
Adaptive guardrails do exist. They are shipping and in production, and I am not going to pretend otherwise. But look at what they adapt to.
Promptfoo markets adaptive guardrails that learn from red-team attack data — probes generated by their own team. That is learning from the people who built the system, not from the world the system operates in. A control that learns only from its builders will discover exactly the failures they were already capable of imagining.
Arthur describes using guardrail failures as input to a self-correction loop. Read it closely: when the guardrail flags a problem, the response is fed back to the model with a correction prompt and the model retries. The guardrail corrects the model. The guardrail’s own criteria never change.
Galileo offers Continuous Learning via Human Feedback, which turns expert reviews into reusable evaluators. That is a human writing new rules, faster.
All three are genuine adaptation. Not one of them closes on the operation.
Even LiSA — the most serious attempt I have found — closes on reported failures about the guardrail’s own allow-or-refuse decisions. A loop on the decision, not on what the decision produced.
One source is honest about the distinction. A thumbs-up button, it says, is a signal, not a loop. A real loop connects feedback to a traceable change — including a guardrail configuration — and then validates that the change improved the outcome.
That is a description of what a loop would require. It is not a description of one that exists.
Road cones
A traffic cone marks where somebody assessed a danger. It does not know whether it is in the right place. It does not learn from being driven past, and it learns nothing from the accident that happens forty metres beyond it.
It sits where somebody put it. Afterward, its only remaining function is to prove that somebody put it there.
That is what a guardrail without a loopback is. Not a control — a reference fallback point, establishing that the rules were followed even when the rules led to failure.
And I want to be precise about the danger, because the overstated version is easy to dismiss.
I am not saying take the cones away. I am saying two narrower things.
A cone that never moves teaches drivers to stop reading cones. You drive past a thousand and touch none. Eventually they stop registering as information about the road and become scenery. By the time one matters, nobody is looking at it.
And the stretch with no cone reads as an all-clear it has not earned. The absence of a guardrail does not mean the road was assessed and found safe. It means nobody assessed it, or assessed it and moved on. Yet an absence is read exactly like a clearance.
They are also wired to the wrong end
Here is the part that troubles me more than the update cycle.
In 1998 I built a system named RAM for a defence maintenance operation, running self-learning algorithms inside what was, by the standards of the day, a nascent AI platform. Its loop closed on the operation. It learned from what happened to the order — delivery outcomes, supplier performance, real conditions in the field. The signal came from the world.
It is worth being concrete about what that meant in practice, because “closes on the operation” is easy to say and easy to fake.
A buyer on the front line could shift the weighting between price and delivery. Not by raising a change request. At the point of the transaction, because they knew something about that order that the standing policy did not.
A manager in the control tower could assess an exception while it was live — read the deviation as it happened and judge its impact, rather than receive it in a report a month later when the only available action was explaining it.
And either of them could introduce a variable the system had no representation for. A major storm closing delivery routes. Nothing in the model knew about the storm. Somebody who could see it put it into the loop.
That last one is the whole distinction. The specification behind that system could do something quite sophisticated on its own: where it could not meet a delivery requirement, it derived an achievable alternate requirement and re-solved against that. A system changing a variable it represents. Impressive, and still bounded — it can only renegotiate the terms it already holds.
The storm is not one of them. The storm arrives from outside the model, through a person who was looking out the window.
That is what a loop closing on the operation means. Not that the system learns faster. That the world has a way in, through people positioned to see what the model does not hold.
Every adaptive guardrail I can find closes on the output. Was the response harmful? Unsupported by the source material? Outside policy? Within scope?
Those are not the same question, and the difference is not academic.
A guardrail can pass every single response in a programme that is failing.
I made this argument in 2008, about SAP’s Gate methodology — a structured checkpoint approach that was, in its day, exactly what better governance was supposed to look like. My assessment then was that it was creative but still worked within the flawed confines of a traditional enterprise-based project. Better discipline applied to an inaccurate representation of the operation is still an inaccurate representation.
Eighteen years on, that is precisely where output-level guardrails sit.
The catch-up game
Now put this in front of autonomous agents.
The literature says the response to a real-world guardrail failure is to collect examples and retrain. Read that as a sequence. The failure occurs. Someone notices. Instances are gathered. The rule is rewritten. The system is updated.
By construction, the guardrail is always one incident behind.
That was survivable when the interval was a quarterly review cycle and a human made every consequential move. It is a different proposition when an agent is acting continuously, at machine speed, on the unfixed condition, for the entire duration of that lag.
And it is worth being clear about what is at the end of that lag. A general at the Department of National Defence put it to me in the early 2000s, and I have carried it since: when a supply chain is disrupted in business, you might wait longer for your pens. In the military, an interruption can cost lives.
That is the distinction that should govern how these controls are built. A control regime can be entirely satisfied while what it exists to protect is failing — and the consequence is not always a variance on a report.
We are running an ineffective control catch-up game. Failures are dissected after the engagement rather than caught at the point of engagement.
That is not a control. It is a post-mortem with a budget line.
What a real loopback would require
This is the uncomfortable part, and it is why nobody has shipped one.
A guardrail that learns needs a signal from the operating reality. Not from the output. Not from the vendor’s red team. From what actually happened when the thing was done.
And that signal does not arrive on its own.
In 1998 the delivery rate sat at fifty-one percent against a ninety percent requirement. The determining condition was that service technicians were batching part orders to late afternoon, because they were measured on call volume and stopping to submit a part order interrupted a service call.
No output check surfaces that. Every response in that system could have been well-formed, within policy, harmless and entirely within scope. The condition that decided the outcome was in a department nobody had put on the map.
So the loopback has the same prerequisite as any other real diagnostic. Somebody has to establish what the operation is actually doing first. Phase 0™ is not an alternative to the control. It is the precondition for building one that can learn.
Delivery reached 97.3% in three months and held for seven years. Cost of goods fell twenty-three percent. Twenty-three full-time equivalents became three inside eighteen months.
Today’s takeaway
Until a guardrail has the real-world, real-time learning attributes that self-learning algorithms had in 1998, it is a road cone. It marks somebody’s assessment of a danger, made in advance, by someone not in the car — and it has no mechanism for finding out whether that assessment was right.
The question to put to any vendor selling AI governance is the one I put to five ProcureTech providers in April 2025, and it has not changed:
Describe the learning loopback. What signal tells your guardrail that it fired on the wrong thing, or that it should have fired and did not, or that the condition it was written for has moved?
If the answer is that your team collects examples and retrains it, that is not a loopback. That is maintenance. And the difference today is that AI moves far faster than past-era technologies, so the interval between the failure and the fix is filled with the system still acting.
And if the answer is that it learns from red-team attacks, ask whose team is running the red team — because a control that learns only from the people who built it will discover exactly the failures they were already capable of imagining.
Keep the human at the wheel. Everything else is just faster.
-30-
This analysis draws on the Procurement Insights archive — an independent record carrying zero vendor sponsorships, published openly since 2007 and consolidating documented client work, lectures, and articles reaching back to 1998. Every claim is held to the Provenance Ledger™: a verify-before-publish discipline that traces each assertion to a primary source and never quietly edits the record once posted. That record is the evidence base for two working lenses — Invariant Physics™, the constant that however far the technology advances the operating logic must be in place first, and Implementation Physics™, its per-engagement application. Phase 0™ identifies and examines the unique and collective attributes within a wide range of seemingly disparate strands — the Strand Commonality™ theory, funded by the Government of Canada’s Scientific Research and Experimental Development program.
Getting it right rather than being right.
Jon W. Hansen, FCIPS — Procurement Insights | Hansen Models™
Related
Your AI Guardrails Are Road Cones
Posted on September 4, 2026
0
A cone does not know whether it is in the right place. It only proves somebody put it there.
Truth Is Believing. Accuracy Is Knowing. Outcome Is Proof.™
Executive Quick Read: Guardrails close on the output; RAM closed on the operation.
Adaptive guardrails exist and are deployed. They adapt to the wrong thing — to attack probes written by the vendor’s own red team, to the model rather than to themselves, or to whatever a human expert last wrote into an evaluator. None of them closes on what happened in the operation.
The research literature says these controls are typically specified before deployment and remain largely fixed during use, and that the response to a real-world failure is to collect examples and retrain. Read as a sequence, that means the guardrail is always one incident behind. Survivable at quarterly review speed. Different when an agent is acting continuously on the unfixed condition.
A guardrail that closes on the output can pass every response in a programme that is failing. In 1998 the determining condition in a defence maintenance operation was a technician incentive in another department — nothing in any output would have shown it.
So the question for anyone selling AI governance is not whether the guardrail adapts. It is what it adapts to.
In April 2025 I looked at how five ProcureTech providers described their agentic AI, and made a general notation:
Eighteen months later, the same test applies to a different subject.
Nobody has described how a guardrail learns.
What the record actually says
This is not a hunch. The field says it about itself.
In May 2026, researchers at Google Cloud AI Research and Seoul National University published LiSA: Lifelong Safety Adaptation via Conservative Policy Induction. Surveying the guardrail field, they find these controls rely on a static, general-purpose definition of harm specified before deployment, and remain largely fixed during use. And they name why that is hard to fix:
The conditions that decide whether a control should fire cannot be listed in advance. That is not my claim about the industry. It is the industry’s claim about itself.
Adaptive guardrails do exist. They are shipping and in production, and I am not going to pretend otherwise. But look at what they adapt to.
Promptfoo markets adaptive guardrails that learn from red-team attack data — probes generated by their own team. That is learning from the people who built the system, not from the world the system operates in. A control that learns only from its builders will discover exactly the failures they were already capable of imagining.
Arthur describes using guardrail failures as input to a self-correction loop. Read it closely: when the guardrail flags a problem, the response is fed back to the model with a correction prompt and the model retries. The guardrail corrects the model. The guardrail’s own criteria never change.
Galileo offers Continuous Learning via Human Feedback, which turns expert reviews into reusable evaluators. That is a human writing new rules, faster.
All three are genuine adaptation. Not one of them closes on the operation.
Even LiSA — the most serious attempt I have found — closes on reported failures about the guardrail’s own allow-or-refuse decisions. A loop on the decision, not on what the decision produced.
One source is honest about the distinction. A thumbs-up button, it says, is a signal, not a loop. A real loop connects feedback to a traceable change — including a guardrail configuration — and then validates that the change improved the outcome.
That is a description of what a loop would require. It is not a description of one that exists.
Road cones
A traffic cone marks where somebody assessed a danger. It does not know whether it is in the right place. It does not learn from being driven past, and it learns nothing from the accident that happens forty metres beyond it.
It sits where somebody put it. Afterward, its only remaining function is to prove that somebody put it there.
That is what a guardrail without a loopback is. Not a control — a reference fallback point, establishing that the rules were followed even when the rules led to failure.
And I want to be precise about the danger, because the overstated version is easy to dismiss.
I am not saying take the cones away. I am saying two narrower things.
A cone that never moves teaches drivers to stop reading cones. You drive past a thousand and touch none. Eventually they stop registering as information about the road and become scenery. By the time one matters, nobody is looking at it.
And the stretch with no cone reads as an all-clear it has not earned. The absence of a guardrail does not mean the road was assessed and found safe. It means nobody assessed it, or assessed it and moved on. Yet an absence is read exactly like a clearance.
They are also wired to the wrong end
Here is the part that troubles me more than the update cycle.
In 1998 I built a system named RAM for a defence maintenance operation, running self-learning algorithms inside what was, by the standards of the day, a nascent AI platform. Its loop closed on the operation. It learned from what happened to the order — delivery outcomes, supplier performance, real conditions in the field. The signal came from the world.
It is worth being concrete about what that meant in practice, because “closes on the operation” is easy to say and easy to fake.
A buyer on the front line could shift the weighting between price and delivery. Not by raising a change request. At the point of the transaction, because they knew something about that order that the standing policy did not.
A manager in the control tower could assess an exception while it was live — read the deviation as it happened and judge its impact, rather than receive it in a report a month later when the only available action was explaining it.
And either of them could introduce a variable the system had no representation for. A major storm closing delivery routes. Nothing in the model knew about the storm. Somebody who could see it put it into the loop.
That last one is the whole distinction. The specification behind that system could do something quite sophisticated on its own: where it could not meet a delivery requirement, it derived an achievable alternate requirement and re-solved against that. A system changing a variable it represents. Impressive, and still bounded — it can only renegotiate the terms it already holds.
The storm is not one of them. The storm arrives from outside the model, through a person who was looking out the window.
That is what a loop closing on the operation means. Not that the system learns faster. That the world has a way in, through people positioned to see what the model does not hold.
Every adaptive guardrail I can find closes on the output. Was the response harmful? Unsupported by the source material? Outside policy? Within scope?
Those are not the same question, and the difference is not academic.
A guardrail can pass every single response in a programme that is failing.
I made this argument in 2008, about SAP’s Gate methodology — a structured checkpoint approach that was, in its day, exactly what better governance was supposed to look like. My assessment then was that it was creative but still worked within the flawed confines of a traditional enterprise-based project. Better discipline applied to an inaccurate representation of the operation is still an inaccurate representation.
Eighteen years on, that is precisely where output-level guardrails sit.
The catch-up game
Now put this in front of autonomous agents.
The literature says the response to a real-world guardrail failure is to collect examples and retrain. Read that as a sequence. The failure occurs. Someone notices. Instances are gathered. The rule is rewritten. The system is updated.
By construction, the guardrail is always one incident behind.
That was survivable when the interval was a quarterly review cycle and a human made every consequential move. It is a different proposition when an agent is acting continuously, at machine speed, on the unfixed condition, for the entire duration of that lag.
And it is worth being clear about what is at the end of that lag. A general at the Department of National Defence put it to me in the early 2000s, and I have carried it since: when a supply chain is disrupted in business, you might wait longer for your pens. In the military, an interruption can cost lives.
That is the distinction that should govern how these controls are built. A control regime can be entirely satisfied while what it exists to protect is failing — and the consequence is not always a variance on a report.
We are running an ineffective control catch-up game. Failures are dissected after the engagement rather than caught at the point of engagement.
That is not a control. It is a post-mortem with a budget line.
What a real loopback would require
This is the uncomfortable part, and it is why nobody has shipped one.
A guardrail that learns needs a signal from the operating reality. Not from the output. Not from the vendor’s red team. From what actually happened when the thing was done.
And that signal does not arrive on its own.
In 1998 the delivery rate sat at fifty-one percent against a ninety percent requirement. The determining condition was that service technicians were batching part orders to late afternoon, because they were measured on call volume and stopping to submit a part order interrupted a service call.
No output check surfaces that. Every response in that system could have been well-formed, within policy, harmless and entirely within scope. The condition that decided the outcome was in a department nobody had put on the map.
So the loopback has the same prerequisite as any other real diagnostic. Somebody has to establish what the operation is actually doing first. Phase 0™ is not an alternative to the control. It is the precondition for building one that can learn.
Delivery reached 97.3% in three months and held for seven years. Cost of goods fell twenty-three percent. Twenty-three full-time equivalents became three inside eighteen months.
Today’s takeaway
Until a guardrail has the real-world, real-time learning attributes that self-learning algorithms had in 1998, it is a road cone. It marks somebody’s assessment of a danger, made in advance, by someone not in the car — and it has no mechanism for finding out whether that assessment was right.
The question to put to any vendor selling AI governance is the one I put to five ProcureTech providers in April 2025, and it has not changed:
Describe the learning loopback. What signal tells your guardrail that it fired on the wrong thing, or that it should have fired and did not, or that the condition it was written for has moved?
If the answer is that your team collects examples and retrains it, that is not a loopback. That is maintenance. And the difference today is that AI moves far faster than past-era technologies, so the interval between the failure and the fix is filled with the system still acting.
And if the answer is that it learns from red-team attacks, ask whose team is running the red team — because a control that learns only from the people who built it will discover exactly the failures they were already capable of imagining.
Keep the human at the wheel. Everything else is just faster.
-30-
This analysis draws on the Procurement Insights archive — an independent record carrying zero vendor sponsorships, published openly since 2007 and consolidating documented client work, lectures, and articles reaching back to 1998. Every claim is held to the Provenance Ledger™: a verify-before-publish discipline that traces each assertion to a primary source and never quietly edits the record once posted. That record is the evidence base for two working lenses — Invariant Physics™, the constant that however far the technology advances the operating logic must be in place first, and Implementation Physics™, its per-engagement application. Phase 0™ identifies and examines the unique and collective attributes within a wide range of seemingly disparate strands — the Strand Commonality™ theory, funded by the Government of Canada’s Scientific Research and Experimental Development program.
Getting it right rather than being right.
Jon W. Hansen, FCIPS — Procurement Insights | Hansen Models™
Share this:
Like this:
Related