Small decision models, compared on the same questions
How small open models compare at one job: reading a text and picking the right answer from a short list.
Last updated 30 September 2026 · results from runs made 27–30 September 2026
A decision model does one narrow job: it reads a text, is asked a question with a fixed list of possible answers — yes or no, one of several categories, a point on a scale — and picks one, with a confidence. That job is most of what automation actually needs: routing a request, flagging spam, judging tone, checking whether a passage answers a question. A large hosted model does it well; a model a hundredth of its size, running on a laptop, may do it well enough, at a fraction of the cost and with nothing leaving the machine. This site measures whether it does. No model wins everywhere — the point is to see, for a given kind of question, what accuracy a given size and speed buys.
The task
One text, one question, a fixed list of answers. Every model is shown exactly the same three things and must pick one answer with a confidence. It never writes free text, which is what makes a 144M encoder and a 35B language model comparable.
How it is scored
Ten public data sets whose right answers were written by people, about 800 questions in all. Every score is shown next to what guessing the most common answer would get, and two models can be compared on the very questions they disagree about, so a gap can be called real or noise.
What is compared
Small encoders built for this job, zero-shot classifiers, small language models fine-tuned to answer with a letter, larger language models used as released, and one hosted service as the reference. Every local model is timed on the same laptop.
Ranking
Every model, ranked by how far above guessing it scores on average across the data sets. Models scored on only some data sets follow, unranked, until they are scored everywhere. Pick another measure or a single data set above the chart; hover a bar for its numbers, click it for the model.
The same ranking as a table, with each model's kind, size and speed. Click a column heading to sort; click a model for its details.
| # | Model | Kind | Size | Runs on | Reads up to | Memory | Gain over guessing | Data sets | Best on | Confidence off by | Per answer |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Jev (hosted) | Hosted | undisclosed | hosted | – | – | +49 pts | all | Yes/no readingIntent routing | 9% | 101 ms |
| 2 | Winnow 12B | Fine-tuned LM | 12B | Apple GPU | 8k | 14 GB | +48 pts | all | Review stars | 11% | 456 ms |
| 3 | Jev-Style 2B | Fine-tuned LM | 2B | Apple GPU | 25k | 6.8 GB | +47 pts | all (5 count) | – | 12% | 55 ms |
| 4 | ruling 35B-A3B | Frozen LM | 35B (3B active) | Apple GPU | 32k | 23 GB | +47 pts | all | – | 12% | 389 ms |
| 5 | Winnow E4B | Fine-tuned LM | E4B | Apple GPU | 8k | 7.7 GB | +47 pts | all | – | 10% | 184 ms |
| 6 | openjev (DiffusionGemma) | Frozen LM | 26B (4B active) | Apple GPU | 32k | 50 GB | +46 pts | all | Financial news | 13% | 386 ms |
| 7 | Jev-Omni 12B | Fine-tuned LM | 12B | Apple GPU | 8k | 29 GB | +45 pts | all | – | 14% | 291 ms |
| 8 | GemmaJev E4B | Frozen LM | E4B | Apple GPU | 4k | 5.4 GB | +44 pts | all | – | 14% | 155 ms |
| 9 | JevK5 4B | Fine-tuned LM | 4B | Apple GPU | 16k | 4.9 GB | +44 pts | all | – | 10% | 180 ms |
| 10 | SemIf 4B | Frozen LM | 4B | Apple GPU | 4k | 10 GB | +43 pts | all | – | 12% | 167 ms |
| 11 | Jev-Style 0.8B | Fine-tuned LM | 0.8B | Apple GPU | 25k | 7.2 GB | +42 pts | all (8 count) | News topicSpamUnfair terms | 12% | 38 ms |
| 12 | Kev 9B · Ollaya | Fine-tuned LM | 9B | CPU | 8k | 26 GB | +42 pts | all (6 count) | – | 14% | 3.1 s |
| 13 | Kev 4B · MLX | Fine-tuned LM | 4B | Apple GPU | 64k | 14 GB | +41 pts | all (6 count) | – | 14% | 124 ms |
| 14 | Kev 4B · Ollaya | Fine-tuned LM | 4B | CPU | 8k | 19 GB | +41 pts | all (6 count) | – | 14% | 1.8 s |
| 15 | decider 2B | Fine-tuned LM | 2B | CPU | 32k | 19 GB | +40 pts | all | – | 8% | 620 ms |
| 16 | Laya multilingual | Decision encoder | 322M | Apple GPU | 1k | 3.6 GB | +37 pts | all | – | 17% | 15 ms |
| 17 | cbjev multilingual | Decision encoder | 322M | Apple GPU | 1k | 3.2 GB | +37 pts | all (5 count) | – | 15% | 19 ms |
| 18 | Laya multilingual · fine-tuned | Decision encoder | 322M | Apple GPU | 1k | 3.6 GB | +36 pts | all | – | 17% | 15 ms |
| 19 | decider 0.8B | Fine-tuned LM | 0.8B | CPU | 32k | 7.5 GB | +36 pts | all | Entailment | 16% | 287 ms |
| 20 | Laya English · Ollaya | Decision encoder | 421M | Apple GPU | 512 | 3.7 GB | +36 pts | all | – | 18% | 31 ms |
| 21 | Decision 1.0 Eos | Fine-tuned LM | 0.8B | CPU | 16k | 8.9 GB | +36 pts | all | – | 11% | 397 ms |
| 22 | Laya multilingual · Ollaya | Decision encoder | 322M | Apple GPU | 1k | 3.3 GB | +36 pts | all | – | 18% | 13 ms |
| 23 | Laya English | Decision encoder | 421M | Apple GPU | 512 | 3.1 GB | +36 pts | all | – | 20% | 36 ms |
| 24 | jeff (GLiFormer) | Zero-shot encoder | 435M | Apple GPU | 1k | 12 GB | +35 pts | all | – | 22% | 90 ms |
| 25 | NLI zero-shot (DeBERTa) | Zero-shot encoder | 435M | CPU | 512 | 3.6 GB | +35 pts | all | – | 14% | 372 ms |
| 26 | jevlike (trained per data set) | Trained per data set | 307M + head | Apple GPU | 512 | 4.1 GB | +30 pts | all | Prompt injection | 27% | 14 ms |
| 27 | Kev 0.8B · Ollaya | Fine-tuned LM | 0.8B | CPU | 8k | 8.4 GB | +29 pts | all (6 count) | – | 14% | 280 ms |
| 28 | Jeff Qwen3.5-2B | Fine-tuned LM | 2B | Apple GPU | 8k | 35 GB | +29 pts | all (6 count) | – | 19% | 45 ms |
| 29 | GLiNER2.5 multi-Decide | Zero-shot encoder | 287M | Apple GPU | 512 | 3.6 GB | +27 pts | all | – | 13% | 35 ms |
| 30 | decider 4B | Fine-tuned LM | 4B | CPU | 32k | 19 GB | +25 pts | all (3 count) | – | 17% | 1.9 s |
| 31 | Von 1.1 | Zero-shot encoder | 395M | CPU | 8k | 10 GB | +24 pts | all | – | 21% | 120 ms |
| 32 | GLiClass instruct large | Zero-shot encoder | 435M | CPU | 1k | 2.9 GB | +20 pts | all | – | 24% | 128 ms |
| 33 | Tacet Sonata | Decision encoder | 144M | Apple GPU | 4k | 3.7 GB | +8 pts | all (9 count) | – | 16% | 12 ms |
| 34 | CLM-8B | Frozen LM | 8B | Apple GPU | 2k | 16 GB | +7 pts | all (1 can't answer) | – | 34% | 112 ms |
| 35 | Multilingual zero-shot (Horizon) | Zero-shot encoder | 308M | Apple GPU | 1k | 1.9 GB | 0 pts | all (3 count) | – | 42% | 21 ms |
| Scored on fewer than 10 data sets — shown, not ranked. An average over fewer sets is not comparable with the ranking above. | |||||||||||
| JEV-9B (AutoTrust) | Fine-tuned LM | 9B | Apple GPU | 1k | 19 GB | +46 pts | 9 of 10 | Finance tweets | 9% | 267 ms | |
| Yardsticks — simple rules, not models. Every model should beat them. | |||||||||||
| Word overlapyardstick | Yardstick | — | CPU | – | 0.0 GB | +2 pts | all | – | 6% | 0 ms | |
| Most common answeryardstick | Yardstick | — | CPU | – | 0.0 GB | −1 pts | all | – | 1% | 0 ms | |
What stands out
- Jev (hosted) has the largest average gain over guessing among the models scored on every data set: +49 pts.
- No single model wins everywhere: 7 different models are the most accurate on at least one data set. Pick by task, not by rank.
- The hardest task is Unfair contract terms (best gain +38 pts); the easiest is Intent routing.
- Of the local models within 10 points of the leader, the fastest is Jev-Style 0.8B at 38 ms per answer with a gain of +42 pts — see speed against gain.
- Some models were trained on the source of a data set they are scored on. Those scores are marked † and left out of that model's average.