Back to Insights
11 Jul 20263 min readAI Agents

Most of your agent's tokens don't need your best model

We watched a week of one agent's traffic and sorted every call by what it was actually doing. The split was not close.

Roughly 70% of the calls were errands: "classify this into one of five buckets," "is this diff risky, yes or no," "extract the three fields," "summarize this log line." The remaining 30% was the real work — the planning, the tricky refactor, the judgment call where being wrong is expensive.

Every single one of those calls went to the same frontier model. The errands and the judgment calls, billed at the same rate.

That is the quiet way an agent bill gets big. Not one giant prompt — a thousand small ones, each reaching for the most expensive model in the drawer to do something a much cheaper model would have done just as well.

The reflex is understandable

You wire the agent to your best model once, at the top, and everything inherits it. It works, so you stop thinking about it. Nobody sits down and decides "this yes/no gate should cost the same as the architecture decision" — it just accretes that way. The default is sticky, and the default is your most capable, most expensive option.

The fix is not "use a cheaper model." It's "use the right model per call," and the only way to do that is to look at the call.

What actually separates the two piles

We stopped guessing and let the task itself decide. A few signals did most of the sorting:

  • Bounded output. If the answer is one of a fixed set — a label, a boolean, a field name — a small model gets it right almost every time. Ambiguity is what needs the big one.
  • Local context. If everything needed to answer is in the prompt, a small model handles it. If the answer requires holding a lot of state and reasoning across it, that's frontier work.
  • Cost of being wrong. A mislabeled log line is noise you'll catch. A wrong migration is a Saturday. Spend the capability where a mistake actually hurts.

None of that requires a model to figure out. A cheap classifier — or even a few rules — can tag a call before it's sent, and route accordingly.

Route, don't switch

The trap on the other side is over-correcting into a routing layer that swaps providers, juggles three SDKs, and quietly changes behavior under you. That trades a cost problem for a reliability problem.

What worked was narrower: stay inside the same provider and the same API, and change only which size of model in that family answers the call. The classifier picks a tier; the request is otherwise untouched. Same auth, same response shape, same failure modes — just not paying frontier rates for a lookup. When we weren't sure a call could be safely downgraded, it defaulted up. Cheaper is the optimization; correct is the constraint.

The part nobody measures

Here's what surprised us: we couldn't answer "where is this agent's money going" until we instrumented it. We had a total bill and no breakdown. Once every call carried a task-type tag, the 70/30 split fell out immediately — and so did the fix.

You cannot right-size what you cannot see. Before any routing, the cheapest win is just a ledger: for each call, how many tokens, which model, what kind of task. Usage, not content — you do not need to log a single message body to learn that most of your spend is errands. Most teams are one afternoon of instrumentation away from finding their own 70%.

The takeaway

Big models are worth every token on the 30% that's genuinely hard. The waste is spending the same on the 70% that isn't. You don't fix that by being cheap; you fix it by looking — tagging each call by what it's doing, and routing the errands to something smaller while the judgment calls keep the best you've got.

The bill isn't big because the work is expensive. It's big because nothing ever asked whether each call needed the expensive answer.

AI agentsLLMscostmodel routingobservability

Found this useful?