Most of your agent's tokens don't need your best model
We watched a week of one agent's traffic and sorted every call by what it was actually doing. The split was not close.
Roughly 70% of the calls were errands: "classify this into one of five buckets," "is this diff risky, yes or no," "extract the three fields," "summarize this log line." The remaining 30% was the real work — the planning, the tricky refactor, the judgment call where being wrong is expensive.
Every single one of those calls went to the same frontier model. The errands and the judgment calls, billed at the same rate.
That is the quiet way an agent bill gets big. Not one giant prompt — a thousand small ones, each reaching for the most expensive model in the drawer to do something a much cheaper model would have done just as well.
The reflex is understandable
You wire the agent to your best model once, at the top, and everything inherits it. It works, so you stop thinking about it. Nobody sits down and decides "this yes/no gate should cost the same as the architecture decision" — it just accretes that way. The default is sticky, and the default is your most capable, most expensive option.
The fix is not "use a cheaper model." It's "use the right model per call," and the only way to do that is to look at the call.
What actually separates the two piles
We stopped guessing and let the task itself decide. A few signals did most of the sorting:
- Bounded output. If the answer is one of a fixed set — a label, a boolean, a field name — a small model gets it right almost every time. Ambiguity is what needs the big one.
- Local context. If everything needed to answer is in the prompt, a small model handles it. If the answer requires holding a lot of state and reasoning across it, that's frontier work.
- Cost of being wrong. A mislabeled log line is noise you'll catch. A wrong migration is a Saturday. Spend the capability where a mistake actually hurts.
None of that requires a model to figure out. A cheap classifier — or even a few rules — can tag a call before it's sent, and route accordingly.
Route, don't switch
The trap on the other side is over-correcting into a routing layer that swaps providers, juggles three SDKs, and quietly changes behavior under you. That trades a cost problem for a reliability problem.
What worked was narrower: stay inside the same provider and the same API, and change only which size of model in that family answers the call. The classifier picks a tier; the request is otherwise untouched. Same auth, same response shape, same failure modes — just not paying frontier rates for a lookup. When we weren't sure a call could be safely downgraded, it defaulted up. Cheaper is the optimization; correct is the constraint.
The part nobody measures
Here's what surprised us: we couldn't answer "where is this agent's money going" until we instrumented it. We had a total bill and no breakdown. Once every call carried a task-type tag, the 70/30 split fell out immediately — and so did the fix.
You cannot right-size what you cannot see. Before any routing, the cheapest win is just a ledger: for each call, how many tokens, which model, what kind of task. Usage, not content — you do not need to log a single message body to learn that most of your spend is errands. Most teams are one afternoon of instrumentation away from finding their own 70%.
The takeaway
Big models are worth every token on the 30% that's genuinely hard. The waste is spending the same on the 70% that isn't. You don't fix that by being cheap; you fix it by looking — tagging each call by what it's doing, and routing the errands to something smaller while the judgment calls keep the best you've got.
The bill isn't big because the work is expensive. It's big because nothing ever asked whether each call needed the expensive answer.
Found this useful?
Related posts
A 16.9 MB model can do the job. Here is how to tell when.
Last week (October 2) Cactus released Whistle, an open speech recognition model in a single 16.9 MB file ([post](https://cactuscompute.com/blog/whistle)). It runs on a CPU with no dependencies, transcribes seven language…
Your agent graded its own homework. Use a second model as the reviewer.
A coding agent finishes a change and reports that it is done and the tests pass. Often both are true, and the change is still wrong: the tests check what the agent thought the code should do, not what you needed.
A small model checking our docs was a coin flip. Giving it the docs fixed most of that.
A Hacker News thread this week argued that agents don't need memory, they need documentation ([discussion](https://news.ycombinator.com/item?id=49945933)). We had just measured a small version of that question, so here i…
We scanned our own agents' Claude Code transcripts for secrets. What we found, and what we changed.
We scanned our coding agents' Claude Code transcripts for secrets, then changed how prompts reach the model. What we found, measured, and what still fails.