A small model checking our docs was a coin flip. Giving it the docs fixed most of that.
A Hacker News thread this week argued that agents don't need memory, they need documentation (discussion). We had just measured a small version of that question, so here is what the numbers said, including the part that didn't go well.
The setup
We keep a wiki where every sentence cites the source lines it came from. When the code changes, some sentences go stale. We wanted a small model, running locally, to flag sentences that are no longer true.
The test set had 148 fresh claims: 74 sentences copied from the current wiki, and 74 copies of those sentences with one number changed. None of them had been used in earlier tests, and nothing was tuned on this set. Five model calls timed out and were dropped, which leaves 143 scored claims.
Asked cold, it was a coin flip
With no context, the model scored an AUC of 0.48 on those 143 claims. That is chance. It wasn't careful, either: it called almost everything false, wrongly failing 67 of the 71 true claims.
The model wasn't too small for the job. It simply had nothing to check the claims against.
Pull the matching docs in first
Before each question, a plain keyword search picks the three wiki sections that best match the claim (about 2,000 characters in all) and puts them in front of the model. No vector database, no embeddings.
With those notes, AUC rose to 0.76 and accuracy to 69%. On paired claims, 47 answers went from wrong to right and 18 from right to wrong (McNemar p = 0.0004).
Add one dumb rule for numbers
The weak spot was changed numbers. So we added a strict rule: if the claim cites a number that doesn't appear in the retrieved sections, fail it.
With memory plus the strict rule, accuracy rose to 81% on the same 143 claims, and 70 of the 72 changed-number claims were caught.
The cost
The same setup also failed 25 of the 71 true claims: about a third.
The rule on its own wrongly flagged 7 of the 74 true sentences (a separate measurement from the 25 above, which is model plus rule), and all 7 had one cause: the cited number really was in the wiki, just not in the three sections the search returned. The rule's false-flag rate equals the retrieval-miss rate. So the next thing to fix is retrieval, not the rule.
It also changed what we tell people. Our first larger run, on 68 claims, showed an AUC of 0.86. At 143 claims it is 0.76. The direction held (p < 0.001); the size shrank. Small test sets flatter you.
What we do with it
The check flags claims for a person. It never decides alone, and it never edits a page by itself. Flagging a third of true claims is fine for a reviewer's queue and bad for anything automatic.
Takeaway
If your agent keeps getting facts wrong, try giving it the right page of your docs before you give it more memory. Then measure on a test set big enough to tell you the truth, and say how big it was.
The claim checker is part of jevx, an open-source command-line tool for small yes/no decisions. deemwar helps teams install, integrate and run it: jevx.deemwar.com.
Found this useful?
Related posts
Six times "done" was wrong, and how to make agent work checkable
Andrej Karpathy wrote on 2 October that we will spend a lot more time trying to understand what language models produce. Teams running coding agents are already paying that cost. The expensive part is rarely the model ge…
You don't need a frontier model for every decision in your agent loop
That distinction is the part worth pulling out, because it maps directly onto a mistake we see in a lot of agent code: every decision point routes through the same big model, because that's the model that's already wired…
Your agent didn't change. The cache did.
Someone ran a fifty-agent pipeline with real cost controls on it: a 200K token ceiling per agent, hard abort checkpoints, runaway guards, a per-agent cost ledger. Thirty-six agents finished 72% of the work on 3–4% of the…
Most of your agent's tokens don't need your best model
We watched a week of one agent's traffic and sorted every call by what it was actually *doing*. The split was not close.