systemone lab

Small decision models, compared on the same questions

How small open models compare at one job: reading a text and picking the right answer from a short list.

A decision model does one narrow job: it reads a text, is asked a question with a fixed list of possible answers — yes or no, one of several categories, a point on a scale — and picks one, with a confidence. That job is most of what automation actually needs: routing a request, flagging spam, judging tone, checking whether a passage answers a question. A large hosted model does it well; a model a hundredth of its size, running on a laptop, may do it well enough, at a fraction of the cost and with nothing leaving the machine. This site measures whether it does. No model wins everywhere — the point is to see, for a given kind of question, what accuracy a given size and speed buys.

36models compared
10public data sets
800questions each model answers
6kinds of model
01

The task

One text, one question, a fixed list of answers. Every model is shown exactly the same three things and must pick one answer with a confidence. It never writes free text, which is what makes a 144M encoder and a 35B language model comparable.

02

How it is scored

Ten public data sets whose right answers were written by people, about 800 questions in all. Every score is shown next to what guessing the most common answer would get, and two models can be compared on the very questions they disagree about, so a gap can be called real or noise.

03

What is compared

Small encoders built for this job, zero-shot classifiers, small language models fine-tuned to answer with a letter, larger language models used as released, and one hosted service as the reference. Every local model is timed on the same laptop.

Ranking

Every model, ranked by how far above guessing it scores on average across the data sets. Models scored on only some data sets follow, unranked, until they are scored everywhere. Pick another measure or a single data set above the chart; hover a bar for its numbers, click it for the model.

The same ranking as a table, with each model's kind, size and speed. Click a column heading to sort; click a model for its details.

#ModelKind SizeRuns on Reads up to Memory Gain over guessing Data sets Best on Confidence off by Per answer
1Jev (hosted)Hostedundisclosedhosted––+49 ptsallYes/no readingIntent routing9%101 ms
2Winnow 12BFine-tuned LM12BApple GPU8k14 GB+48 ptsallReview stars11%456 ms
3Jev-Style 2BFine-tuned LM2BApple GPU25k6.8 GB+47 ptsall (5 count)–12%55 ms
4ruling 35B-A3BFrozen LM35B (3B active)Apple GPU32k23 GB+47 ptsall–12%389 ms
5Winnow E4BFine-tuned LME4BApple GPU8k7.7 GB+47 ptsall–10%184 ms
6openjev (DiffusionGemma)Frozen LM26B (4B active)Apple GPU32k50 GB+46 ptsallFinancial news13%386 ms
7Jev-Omni 12BFine-tuned LM12BApple GPU8k29 GB+45 ptsall–14%291 ms
8GemmaJev E4BFrozen LME4BApple GPU4k5.4 GB+44 ptsall–14%155 ms
9JevK5 4BFine-tuned LM4BApple GPU16k4.9 GB+44 ptsall–10%180 ms
10SemIf 4BFrozen LM4BApple GPU4k10 GB+43 ptsall–12%167 ms
11Jev-Style 0.8BFine-tuned LM0.8BApple GPU25k7.2 GB+42 ptsall (8 count)News topicSpamUnfair terms12%38 ms
12Kev 9B · OllayaFine-tuned LM9BCPU8k26 GB+42 ptsall (6 count)–14%3.1 s
13Kev 4B · MLXFine-tuned LM4BApple GPU64k14 GB+41 ptsall (6 count)–14%124 ms
14Kev 4B · OllayaFine-tuned LM4BCPU8k19 GB+41 ptsall (6 count)–14%1.8 s
15decider 2BFine-tuned LM2BCPU32k19 GB+40 ptsall–8%620 ms
16Laya multilingualDecision encoder322MApple GPU1k3.6 GB+37 ptsall–17%15 ms
17cbjev multilingualDecision encoder322MApple GPU1k3.2 GB+37 ptsall (5 count)–15%19 ms
18Laya multilingual · fine-tunedDecision encoder322MApple GPU1k3.6 GB+36 ptsall–17%15 ms
19decider 0.8BFine-tuned LM0.8BCPU32k7.5 GB+36 ptsallEntailment16%287 ms
20Laya English · OllayaDecision encoder421MApple GPU5123.7 GB+36 ptsall–18%31 ms
21Decision 1.0 EosFine-tuned LM0.8BCPU16k8.9 GB+36 ptsall–11%397 ms
22Laya multilingual · OllayaDecision encoder322MApple GPU1k3.3 GB+36 ptsall–18%13 ms
23Laya EnglishDecision encoder421MApple GPU5123.1 GB+36 ptsall–20%36 ms
24jeff (GLiFormer)Zero-shot encoder435MApple GPU1k12 GB+35 ptsall–22%90 ms
25NLI zero-shot (DeBERTa)Zero-shot encoder435MCPU5123.6 GB+35 ptsall–14%372 ms
26jevlike (trained per data set)Trained per data set307M + headApple GPU5124.1 GB+30 ptsallPrompt injection27%14 ms
27Kev 0.8B · OllayaFine-tuned LM0.8BCPU8k8.4 GB+29 ptsall (6 count)–14%280 ms
28Jeff Qwen3.5-2BFine-tuned LM2BApple GPU8k35 GB+29 ptsall (6 count)–19%45 ms
29GLiNER2.5 multi-DecideZero-shot encoder287MApple GPU5123.6 GB+27 ptsall–13%35 ms
30decider 4BFine-tuned LM4BCPU32k19 GB+25 ptsall (3 count)–17%1.9 s
31Von 1.1Zero-shot encoder395MCPU8k10 GB+24 ptsall–21%120 ms
32GLiClass instruct largeZero-shot encoder435MCPU1k2.9 GB+20 ptsall–24%128 ms
33Tacet SonataDecision encoder144MApple GPU4k3.7 GB+8 ptsall (9 count)–16%12 ms
34CLM-8BFrozen LM8BApple GPU2k16 GB+7 ptsall (1 can't answer)–34%112 ms
35Multilingual zero-shot (Horizon)Zero-shot encoder308MApple GPU1k1.9 GB0 ptsall (3 count)–42%21 ms
Scored on fewer than 10 data sets — shown, not ranked. An average over fewer sets is not comparable with the ranking above.
JEV-9B (AutoTrust)Fine-tuned LM9BApple GPU1k19 GB+46 pts9 of 10Finance tweets9%267 ms
Yardsticks — simple rules, not models. Every model should beat them.
Word overlapyardstickYardstick—CPU–0.0 GB+2 ptsall–6%0 ms
Most common answeryardstickYardstick—CPU–0.0 GB−1 ptsall–1%0 ms

What stands out