Small decision models, compared on the same questions
How small open models compare at one job: reading a text and picking the right answer from a short list.
Last updated 29 September 2026 · results from runs made 27–29 September 2026
A decision model does one narrow job: it reads a text, is asked a question with a fixed list of possible answers — yes or no, one of several categories, a point on a scale — and picks one, with a confidence. That job is most of what automation actually needs: routing a request, flagging spam, judging tone, checking whether a passage answers a question. A large hosted model does it well; a model a hundredth of its size, running on a laptop, may do it well enough, at a fraction of the cost and with nothing leaving the machine. This site measures whether it does. No model wins everywhere — the point is to see, for a given kind of question, what accuracy a given size and speed buys.
The task
One text, one question, a fixed list of answers. Every model is shown exactly the same three things and must pick one answer with a confidence. It never writes free text, which is what makes a 144M encoder and a 35B language model comparable.
How it is scored
Ten public data sets whose right answers were written by people, about 800 questions in all. Every score is shown next to what guessing the most common answer would get, and two models can be compared on the very questions they disagree about, so a gap can be called real or noise.
What is compared
Small encoders built for this job, zero-shot classifiers, small language models fine-tuned to answer with a letter, larger language models used as released, and one hosted service as the reference. Every local model is timed on the same laptop.
Ranking
Every model, ranked by how far above guessing it scores on average across the data sets. Click a column heading to sort; click a model for its details.
| # | Model | Kind | Size | Runs on | Gain over guessing | Data sets | Best on | Confidence off by | Per answer |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Jev (hosted) | Hosted | undisclosed | hosted | +49 pts | all | Review starsYes/no readingIntent routing | 9% | 101 ms |
| 2 | ruling 35B-A3B | Frozen LM | 35B (3B active) | Apple GPU | +47 pts | all | Finance tweets | 12% | 391 ms |
| 3 | Winnow E4B | Fine-tuned LM | E4B | Apple GPU | +47 pts | all | Financial news | 10% | 184 ms |
| 4 | JevK5 4B | Fine-tuned LM | 4B | Apple GPU | +44 pts | all | – | 10% | 181 ms |
| 5 | SemIf 4B | Frozen LM | 4B | Apple GPU | +43 pts | all | – | 12% | 171 ms |
| 6 | Jev-Style 0.8B | Fine-tuned LM | 0.8B | Apple GPU | +42 pts | all (8 count) | News topicSpamUnfair terms | 12% | 39 ms |
| 7 | Kev 9B · Ollaya | Fine-tuned LM | 9B | CPU | +42 pts | all (6 count) | – | 14% | 3.2 s |
| 8 | Kev 4B · MLX | Fine-tuned LM | 4B | Apple GPU | +41 pts | all (6 count) | – | 14% | 131 ms |
| 9 | Kev 4B · Ollaya | Fine-tuned LM | 4B | CPU | +41 pts | all (6 count) | – | 14% | 1.8 s |
| 10 | decider 2B | Fine-tuned LM | 2B | CPU | +40 pts | all | – | 8% | 614 ms |
| 11 | Laya multilingual | Decision encoder | 322M | Apple GPU | +37 pts | all | – | 17% | 15 ms |
| 12 | cbjev multilingual | Decision encoder | 322M | Apple GPU | +37 pts | all (5 count) | – | 15% | 20 ms |
| 13 | Laya multilingual · fine-tuned | Decision encoder | 322M | Apple GPU | +36 pts | all | – | 17% | 15 ms |
| 14 | decider 0.8B | Fine-tuned LM | 0.8B | CPU | +36 pts | all | Entailment | 16% | 279 ms |
| 15 | Laya English · Ollaya | Decision encoder | 421M | Apple GPU | +36 pts | all | – | 18% | 32 ms |
| 16 | Decision 1.0 Eos | Fine-tuned LM | 0.8B | CPU | +36 pts | all | – | 11% | 391 ms |
| 17 | Laya multilingual · Ollaya | Decision encoder | 322M | Apple GPU | +36 pts | all | – | 18% | 13 ms |
| 18 | Laya English | Decision encoder | 421M | Apple GPU | +36 pts | all | – | 20% | 35 ms |
| 19 | jeff (GLiFormer) | Zero-shot encoder | 435M | Apple GPU | +35 pts | all | – | 22% | 90 ms |
| 20 | NLI zero-shot (DeBERTa) | Zero-shot encoder | 435M | CPU | +35 pts | all | – | 14% | 366 ms |
| 21 | jevlike (trained per data set) | Trained per data set | 307M + head | Apple GPU | +30 pts | all | Prompt injection | 27% | 14 ms |
| 22 | Kev 0.8B · Ollaya | Fine-tuned LM | 0.8B | CPU | +29 pts | all (6 count) | – | 14% | 286 ms |
| 23 | Von 1.1 | Zero-shot encoder | 395M | CPU | +24 pts | all | – | 21% | 120 ms |
| 24 | GLiClass instruct large | Zero-shot encoder | 435M | CPU | +20 pts | all | – | 24% | 123 ms |
| 25 | Tacet Sonata | Decision encoder | 144M | Apple GPU | +8 pts | all (9 count) | – | 16% | 13 ms |
| 26 | CLM-8B | Frozen LM | 8B | Apple GPU | +7 pts | all (1 can't answer) | – | 34% | 114 ms |
| Yardsticks — simple rules, not models. Every model should beat them. | |||||||||
| Word overlapyardstick | Yardstick | — | CPU | +2 pts | all | – | 6% | 0 ms | |
| Most common answeryardstick | Yardstick | — | CPU | −1 pts | all | – | 1% | 0 ms | |
What stands out
- Jev (hosted) has the largest average gain over guessing among the models scored on every data set: +49 pts.
- No single model wins everywhere: 6 different models are the most accurate on at least one data set. Pick by task, not by rank.
- The hardest task is Unfair contract terms (best gain +38 pts); the easiest is Intent routing.
- Of the local models within 10 points of the leader, the fastest is Jev-Style 0.8B at 39 ms per answer with a gain of +42 pts — see speed against gain.
- Some models were trained on the source of a data set they are scored on. Those scores are marked † and left out of that model's average.