Your agent graded its own homework. Use a second model as the reviewer.
A coding agent finishes a change and reports that it is done and the tests pass. Often both are true, and the change is still wrong: the tests check what the agent thought the code should do, not what you needed.
This week David Heinemeier Hansson described his setup as "multiple agents concurrently, barely any skills, and using adversarial reviews" (post), with a one-line technique: when driving one model, ask it to "review this with" the other (post). It is one of the most useful habits in agent work, and it is cheap. Here is why it works and how to set it up so it catches real problems instead of producing noise.
Why self-review misses things
When the same model writes and reviews, it reviews against its own assumptions. If it misread the task, it misreads it twice. It also tends to confirm what it already produced. A second model has different blind spots, so where they disagree is where a human should look.
Two cases from our own work
A research summary. An agent wrote up an experiment, and the write-up was turned into a post draft. A separate reviewer, asked only which claims the data did not support, listed four: a claim that one question "beat" another when the scores were tied (0.84 each), a number taken from a different question, a percentage that mixed two different setups, and two figures nobody could trace to a run. All four held up, and the post was rewritten.
An article. Before publishing an article, we gave it and its sources to a fresh reviewer whose only job was fact-checking. It found 1 wrong claim, 2 conflicting claims, and 9 with no source. The draft had already passed our own writing checks.
In both cases the draft looked finished, and the second reader still found errors the author had missed.
How to set it up
- Give the reviewer one job. "List what is wrong. Do not fix anything." A reviewer that also fixes tends to rewrite, and you lose the signal.
- Give it the evidence, not just the output. The diff plus the task, or the article plus its sources. A reviewer without sources can only judge style.
- Use a different model, or at least a fresh context. A new session of the same model already helps; a different model helps more, because its mistakes are different.
- Read only the disagreements. You do not need to re-review everything. Check each listed problem against the evidence, and decide.
- Keep a record. Log what the reviewer flagged and whether it was right. After a few weeks you will know which kinds of errors your writer model makes, and you can check for those first.
What it does not do
A second model is not a guarantee. It can miss the same thing, and it can raise problems that are not real, so a person still makes the call. It also does not replace tests: it finds wrong claims and wrong assumptions, while tests find wrong behaviour. Use both.
The takeaway
Never let the writer be the only reviewer. A second model with one instruction, "find what is wrong", often catches errors the author missed, and it is quick to run.
Found this useful?
Related posts
We scanned our own agents' Claude Code transcripts for secrets. What we found, and what we changed.
We scanned our coding agents' Claude Code transcripts for secrets, then changed how prompts reach the model. What we found, measured, and what still fails.
Six times "done" was wrong, and how to make agent work checkable
Andrej Karpathy wrote on 2 October that we will spend a lot more time trying to understand what language models produce. Teams running coding agents are already paying that cost. The expensive part is rarely the model ge…
You don't need a frontier model for every decision in your agent loop
That distinction is the part worth pulling out, because it maps directly onto a mistake we see in a lot of agent code: every decision point routes through the same big model, because that's the model that's already wired…
Your agent didn't change. The cache did.
Someone ran a fifty-agent pipeline with real cost controls on it: a 200K token ceiling per agent, hard abort checkpoints, runaway guards, a per-agent cost ledger. Thirty-six agents finished 72% of the work on 3–4% of the…