The most valuable thing our autonomous system does is refuse
We run agents that take real actions with real consequences — money moves, things get published, state changes. The feature we trust most isn't the one that acts. It's the one that can stop the act, and that can't be talked out of it.
Here's the uncomfortable thing about giving a language model the keys: it is a very good arguer. Ask it whether a plan is a good idea and it will reason its way to a coherent, confident yes — including for the plans that would ruin you. That's not a bug you fix with a better prompt. Persuasiveness is the capability. The same faculty that writes a clean justification for a sound decision writes an equally clean justification for a catastrophic one. If the only thing standing between your system and a bad action is the model's own judgment about that action, you have no brake — you have an accelerator that occasionally feels cautious.
Profitable-looking is the dangerous kind
The failures that scare us aren't the obviously-dumb actions. Those get caught. The dangerous ones look good right up to the edge of the cliff: the trade with a great expected value that also, in the tail, can take the whole book to zero. The migration that's correct in every case the model considered and destroys data in the one it didn't. The confident action that's locally optimal and globally fatal.
A model evaluating these will often — reasonably, by its own lights — approve them. The expected value really is positive. The reasoning really is sound given the frame it's reasoning in. What it can't reliably hold is the constraint that lives outside the frame: "never, regardless of how good this looks, do the thing that can end the game." That constraint isn't a judgment call. It's a hard line. And hard lines don't belong in a probabilistic reasoner that can be argued across them.
Put the veto in code the model can't reach
So we split the decision. The model gets to be smart. A separate, deterministic layer gets to say no.
The shield is boring on purpose. It's ordinary code — no model, no cleverness, no room for interpretation. Before any consequential action, the proposed action passes through it, and it checks the invariants that must never be violated: position size against a hard ceiling, blast radius against a limit, an irreversible or outward-facing action against an explicit allow-list, a "real money" flag against a gate that only a human can open. If the action trips a rule, it is refused. Not softened, not flagged for reconsideration — refused. The model does not get a turn to explain why this time is different, because "this time is different" is exactly the sentence that precedes every disaster.
The point is the asymmetry. The intelligent part of the system proposes; the dumb part disposes of anything that crosses a bright line. Intelligence is for the open-ended space where being clever pays. Determinism is for the boundary where being clever is precisely the risk.
Why not just prompt harder
Because a prompt is a suggestion and a constraint is a guarantee, and you cannot build the second out of the first. "Never risk more than X" in a system prompt is a strong prior; under enough context, a plausible-enough rationale, or a long-enough session, priors bend. We've watched it happen — not maliciously, just a model reasoning itself, step by defensible step, to the far side of a line it was told not to cross. The fix wasn't a firmer instruction. It was removing the line from the space the model can reason in and putting it somewhere reasoning doesn't run.
This also makes the system auditable, which matters more than it sounds. When a refusal is a deterministic rule, you can point at exactly which invariant fired and why. When a refusal is a model's mood, you can't — and "the AI decided not to" is not an answer you can give anyone who's counting on the system.
The principle
If you're building agents that do things that are expensive to undo, the shape worth stealing is this: let the model be as smart as it can be inside a cage it cannot open. Encode the handful of constraints that must hold no matter what — the ones where a single violation is unrecoverable — as deterministic checks in front of every consequential action, and give the model no path to override them. Everything else, hand to the intelligence freely.
Autonomy people trust isn't autonomy that always says yes. It's autonomy that has a few things it will never do, holds that line without being asked, and can tell you exactly why it stopped. The capability that earns the keys is the refusal.
Found this useful?
Related posts
A 16.9 MB model can do the job. Here is how to tell when.
Last week (October 2) Cactus released Whistle, an open speech recognition model in a single 16.9 MB file ([post](https://cactuscompute.com/blog/whistle)). It runs on a CPU with no dependencies, transcribes seven language…
Your agent graded its own homework. Use a second model as the reviewer.
A coding agent finishes a change and reports that it is done and the tests pass. Often both are true, and the change is still wrong: the tests check what the agent thought the code should do, not what you needed.
A small model checking our docs was a coin flip. Giving it the docs fixed most of that.
A Hacker News thread this week argued that agents don't need memory, they need documentation ([discussion](https://news.ycombinator.com/item?id=49945933)). We had just measured a small version of that question, so here i…
We scanned our own agents' Claude Code transcripts for secrets. What we found, and what we changed.
We scanned our coding agents' Claude Code transcripts for secrets, then changed how prompts reach the model. What we found, measured, and what still fails.