Back to Insights
5 Oct 20263 min readAI Agents

A small model checking our docs was a coin flip. Giving it the docs fixed most of that.

A Hacker News thread this week argued that agents don't need memory, they need documentation (discussion). We had just measured a small version of that question, so here is what the numbers said, including the part that didn't go well.

The setup

We keep a wiki where every sentence cites the source lines it came from. When the code changes, some sentences go stale. We wanted a small model, running locally, to flag sentences that are no longer true.

The test set had 148 fresh claims: 74 sentences copied from the current wiki, and 74 copies of those sentences with one number changed. None of them had been used in earlier tests, and nothing was tuned on this set. Five model calls timed out and were dropped, which leaves 143 scored claims.

Asked cold, it was a coin flip

With no context, the model scored an AUC of 0.48 on those 143 claims. That is chance. It wasn't careful, either: it called almost everything false, wrongly failing 67 of the 71 true claims.

The model wasn't too small for the job. It simply had nothing to check the claims against.

Pull the matching docs in first

Before each question, a plain keyword search picks the three wiki sections that best match the claim (about 2,000 characters in all) and puts them in front of the model. No vector database, no embeddings.

With those notes, AUC rose to 0.76 and accuracy to 69%. On paired claims, 47 answers went from wrong to right and 18 from right to wrong (McNemar p = 0.0004).

Add one dumb rule for numbers

The weak spot was changed numbers. So we added a strict rule: if the claim cites a number that doesn't appear in the retrieved sections, fail it.

With memory plus the strict rule, accuracy rose to 81% on the same 143 claims, and 70 of the 72 changed-number claims were caught.

The cost

The same setup also failed 25 of the 71 true claims: about a third.

The rule on its own wrongly flagged 7 of the 74 true sentences (a separate measurement from the 25 above, which is model plus rule), and all 7 had one cause: the cited number really was in the wiki, just not in the three sections the search returned. The rule's false-flag rate equals the retrieval-miss rate. So the next thing to fix is retrieval, not the rule.

It also changed what we tell people. Our first larger run, on 68 claims, showed an AUC of 0.86. At 143 claims it is 0.76. The direction held (p < 0.001); the size shrank. Small test sets flatter you.

What we do with it

The check flags claims for a person. It never decides alone, and it never edits a page by itself. Flagging a third of true claims is fine for a reviewer's queue and bad for anything automatic.

Takeaway

If your agent keeps getting facts wrong, try giving it the right page of your docs before you give it more memory. Then measure on a test set big enough to tell you the truth, and say how big it was.


The claim checker is part of jevx, an open-source command-line tool for small yes/no decisions. deemwar helps teams install, integrate and run it: jevx.deemwar.com.

AI agentsevaluationdocumentation