Back to Insights
9 Oct 20263 min readAI Agents

A 16.9 MB model can do the job. Here is how to tell when.

Last week (October 2) Cactus released Whistle, an open speech recognition model in a single 16.9 MB file (post). It runs on a CPU with no dependencies, transcribes seven languages, handles clips of up to 30 seconds, and reports its first token in 11 ms. It is built for phones, wearables, cars and microcontrollers, where a large hosted model is not an option.

It is a good reminder that many jobs in a product do not need the biggest model. They need a small one that does one narrow thing, close to the data. The hard part is knowing when small is enough.

Small models do well on narrow, checkable jobs

A narrow job has a fixed input, a short answer and a clear way to grade it: transcribe this clip, is this message a prompt injection, does this sentence match the docs. These are the jobs where a small model on your own hardware pays off: no network hop, no data leaving your machine, and a cost that does not grow with every call.

Two results from our own testing show both sides.

Prompt injection. Our keyword filter missed most of the injections in a public test set, because each one was worded differently. A small local model running on a CPU flagged 70% of the injections in a public test set. That is not enough on its own, so it now runs alongside the filter, not instead of it.

Checking claims against docs. Asked cold, a small local model was no better than a coin flip at spotting wrong claims in documentation. With the matching doc sections pulled in first, it did much better on 143 fresh claims, and in its strict mode it caught 70 of 72 changed numbers. It also rejected about a third of true claims, so it flags for a person and never decides alone.

The second case is the important one. The same small model went from useless to useful, and the model did not change. The input did.

How to tell if a small model is enough

  1. Write down the one question it answers. If you cannot state it in a sentence, the job is too broad for a small model.
  2. Build a test set from your own data, with the answers. A public benchmark tells you little about your inputs. A few hundred labelled examples from your own traffic tell you a lot.
  3. Measure both kinds of error. A spam filter that blocks real customers is worse than one that lets some spam through. Decide which mistake costs more before you pick a threshold.
  4. Give it context before you give up on it. A small model asked cold often fails where the same model with the right few paragraphs succeeds.
  5. Decide what happens when it is unsure. Send those cases to a bigger model or a person, and count how often that happens.

What small models still do not do

They do not reason well over long, open-ended tasks, and they degrade quickly outside the narrow job they were trained or tested for. Whistle, for example, covers seven languages; a clip in an eighth is outside its job. Treat a small model as a component with a known accuracy on a known input, not as a general assistant.

The takeaway

Start with the narrowest question, measure on your own data, and add context before you reach for a bigger model. When the job is narrow and checkable, a small model close to the data is often the better engineering choice, not the cheaper compromise.

AI agentssmall modelsevaluation

Found this useful?