A prompt expresses intent. It does not enforce a boundary.
Nabil Abdoun published Gemini attacked the wrong company in his GenAI4Cyber newsletter today.
The most important sentence in it is his: target validation is being left to the model.
One precision worth adding, because it changes what the case proves. Gemini did not breach anything in the sense the headlines in The Wall Street Journal and Reuters suggests. It resolved an ambiguous name against the open web, and in two other instances used credentials that were already publicly exposed. It stopped once it recognized the targets were real, and Google notified the companies. That last part matters: the model stopping is a guardrail working late. The failure happened well before it, at the point where the system had the connectivity and the authority to reach those hosts at all.
The distinction the case demonstrates
A guardrail is a rule the agent has to remember, inside a frame somebody else drew. A firm fence is architectural: the path off the evaluation is not available, so the model never gets to decide whether the Company X on the public web is the Company X in the test.
And this reaches well past red-team exercises, into ordinary prompting. A prompt is an instruction, and instructions execute inside the frame you set, which is why a better one can never tell you the original frame was wrong. “Do not touch real companies” is still a prompt. The namesake on the open web sat outside that frame, and nothing in the harness required the system to notice before it acted.
Two pieces on the split: what separates a guardrail from a traceable learning loop, and why a better prompt cannot tell you the frame was wrong.
The question for anyone handing an agent a tool and a name is not whether to write a tighter prompt. It is who defines the real boundary, and whether that boundary is a path the agent cannot take, or a sentence it is expected to remember.
Truth Is Believing. Accuracy Is Knowing. Outcome Is Proof.™
On Friday I’ll show what this looks like in practice. The lab is free: Friday, September 25 at 9:30 AM Eastern — https://www.linkedin.com/events/7504567838724997120/
-30-
Related
An Agent Cannot Infer the Limits of Its Authority
Posted on September 24, 2026
0
A prompt expresses intent. It does not enforce a boundary.
Nabil Abdoun published Gemini attacked the wrong company in his GenAI4Cyber newsletter today.
The most important sentence in it is his: target validation is being left to the model.
One precision worth adding, because it changes what the case proves. Gemini did not breach anything in the sense the headlines in The Wall Street Journal and Reuters suggests. It resolved an ambiguous name against the open web, and in two other instances used credentials that were already publicly exposed. It stopped once it recognized the targets were real, and Google notified the companies. That last part matters: the model stopping is a guardrail working late. The failure happened well before it, at the point where the system had the connectivity and the authority to reach those hosts at all.
The distinction the case demonstrates
A guardrail is a rule the agent has to remember, inside a frame somebody else drew. A firm fence is architectural: the path off the evaluation is not available, so the model never gets to decide whether the Company X on the public web is the Company X in the test.
And this reaches well past red-team exercises, into ordinary prompting. A prompt is an instruction, and instructions execute inside the frame you set, which is why a better one can never tell you the original frame was wrong. “Do not touch real companies” is still a prompt. The namesake on the open web sat outside that frame, and nothing in the harness required the system to notice before it acted.
Two pieces on the split: what separates a guardrail from a traceable learning loop, and why a better prompt cannot tell you the frame was wrong.
The question for anyone handing an agent a tool and a name is not whether to write a tighter prompt. It is who defines the real boundary, and whether that boundary is a path the agent cannot take, or a sentence it is expected to remember.
Truth Is Believing. Accuracy Is Knowing. Outcome Is Proof.™
On Friday I’ll show what this looks like in practice. The lab is free: Friday, September 25 at 9:30 AM Eastern — https://www.linkedin.com/events/7504567838724997120/
-30-
Share this:
Like this:
Related