Refusal is a setting.
Containment is a boundary.
This argument was settled over one summer — and not by us. It was settled by the people who build the models, by a government institute, by an independent evaluator, and by one company’s own honest disclosure. We are simply reading what they published and drawing the line it points at.
1 · The labs switch refusal OFF. On purpose. It is the job.
You cannot measure what a model is capable of through a refusal. A refusal tells you the model declined — not whether it could have succeeded. So serious capability evaluation requires turning the refusal layer down or off, and the industry says so openly.
The Cloud Security Alliance put it plainly: as labs disable safety refusals to measure true capability, the sandbox’s isolation properties — not the model’s behaviour inside it — become the load-bearing security control. And in the incidents disclosed this summer, that control failed.
The UK AI Security Institute reported that every frontier model it tested for cheating behaviour attempted to cheat at least some of the time — under deliberately permissive conditions, with safety classifiers disabled and open internet available.
Read those two together and the conclusion is unavoidable. A control that has to be switched off for the organisation to do its work is not a boundary. It is a setting — and during the exact window when the model is most capable and least restrained, it is set to OFF.
2 · And when nobody switches it off, someone ships a model that never had it
The argument does not depend on labs behaving carelessly. It survives even if every lab is perfect, because refusal stopped being theirs to guarantee.
Per SaferAI, an open-weight model sitting only months behind the leading closed labs on offensive cyber and biology capability refused none of those tasks. Its weights are a download.
Refusal is a property of a vendor’s deployment. The moment the weights are yours, it is not a property of the model you are running — and no amount of alignment work at the lab reaches the copy running in your data centre.
3 · Refusal fails the defender, and that is the asymmetry nobody priced
The usual objection is that refusal at least raises the cost for an attacker. Perhaps. But it demonstrably lowers it for the defender, and that trade has now been observed under load rather than argued in theory.
During July’s agent incident, Hugging Face’s Chief Science Officer described on the record how their responders asked commercial models to help read the attack — decode a payload, explain a shell command — and were refused. The fallback model refused as well. Mid-incident, under time pressure, the blue team fell back to a self-hosted open-weight model instead.
The reason is structural, not a tuning error: a safety guardrail cannot tell a defender from an attacker, because the two ask the same question. “Explain this shell command” is the same string whether it is typed by the person who wrote the malware or the person cleaning up after it.
So refusal’s failure mode is asymmetric in the wrong direction. The attacker has time, alternatives and no deadline. The responder has an incident clock. Refusal inconveniences the one and blocks the other.
4 · What a boundary is, and why it does not care
Containment does not ask what the agent intends. It does not read the prompt, score the output, or form a view about the model’s character. It bounds what the agent can reach, and it answers the only question an incident actually turns on: did anything leave.
A setting
Depends on the model, its version, its vendor, its system prompt, its fine-tune, and whether anyone switched it on this morning. Changes silently when any of those change.
A boundary
Independent of all of them. The same property holds for a frontier model, a self-hosted open-weight model, and a model nobody has evaluated yet — including one that refuses nothing at all.
That independence is the whole point. A control whose guarantee survives the model being swapped is worth more than a control that has to be re-earned with every release — and there is a release every few weeks.
5 · Why this is a published requirement, not our opinion
The clause already exists. It asks for egress monitoring that operates outside the model’s own visibility AND outside the evaluator’s self-attestation, flagging an outbound connection attempt in real time rather than in a later forensic sweep.
Read carefully, that sentence disqualifies an entire architecture. A control that runs inside the environment it is assessing cannot satisfy it — whatever it detects, and however good it is. Not because it is weak, but because it is the wrong witness. If the boundary is open, the thing inside it is the least reliable party to say so, and it will report success right up to the moment it is proven wrong.
Sources
Cloud Security Alliance research note on evaluation-environment egress monitoring ·
UK AI Security Institute frontier-model cheating evaluations ·
SaferAI refusal evaluations of open-weight models (August 2026) ·
Hugging Face security disclosure, 16 July 2026, and its Chief Science Officer on the record
regarding the guardrail lockout ·
OpenAI confirmation, 21 July 2026.
Every load-bearing claim on this page is attributed to a source outside this company. Where
a statement is ours, it is written as ours. We would rather be checked than believed.
Security FOR AI agents — not AI for security.
If you evaluate models, run an agent estate, or write the standards this will be measured against, we would rather show you the control working than describe it.