systemone lab

Small decision models, compared on the same questions

How small open models compare at one job: reading a text and picking the right answer from a short list.

A decision model does one narrow job: it reads a text, is asked a question with a fixed list of possible answers — yes or no, one of several categories, a point on a scale — and picks one, with a confidence. That job is most of what automation actually needs: routing a request, flagging spam, judging tone, checking whether a passage answers a question. A large hosted model does it well; a model a hundredth of its size, running on a laptop, may do it well enough, at a fraction of the cost and with nothing leaving the machine. This site measures whether it does. No model wins everywhere — the point is to see, for a given kind of question, what accuracy a given size and speed buys.

26models compared
10public data sets
800questions each model answers
6kinds of model
01

The task

One text, one question, a fixed list of answers. Every model is shown exactly the same three things and must pick one answer with a confidence. It never writes free text, which is what makes a 144M encoder and a 35B language model comparable.

02

How it is scored

Ten public data sets whose right answers were written by people, about 800 questions in all. Every score is shown next to what guessing the most common answer would get, and two models can be compared on the very questions they disagree about, so a gap can be called real or noise.

03

What is compared

Small encoders built for this job, zero-shot classifiers, small language models fine-tuned to answer with a letter, larger language models used as released, and one hosted service as the reference. Every local model is timed on the same laptop.

Ranking

Every model, ranked by how far above guessing it scores on average across the data sets. Click a column heading to sort; click a model for its details.

#ModelKind SizeRuns on Gain over guessing Data sets Best on Confidence off by Per answer
1Jev (hosted)Hostedundisclosedhosted+49 ptsallReview starsYes/no readingIntent routing9%101 ms
2ruling 35B-A3BFrozen LM35B (3B active)Apple GPU+47 ptsallFinance tweets12%391 ms
3Winnow E4BFine-tuned LME4BApple GPU+47 ptsallFinancial news10%184 ms
4JevK5 4BFine-tuned LM4BApple GPU+44 ptsall–10%181 ms
5SemIf 4BFrozen LM4BApple GPU+43 ptsall–12%171 ms
6Jev-Style 0.8BFine-tuned LM0.8BApple GPU+42 ptsall (8 count)News topicSpamUnfair terms12%39 ms
7Kev 9B · OllayaFine-tuned LM9BCPU+42 ptsall (6 count)–14%3.2 s
8Kev 4B · MLXFine-tuned LM4BApple GPU+41 ptsall (6 count)–14%131 ms
9Kev 4B · OllayaFine-tuned LM4BCPU+41 ptsall (6 count)–14%1.8 s
10decider 2BFine-tuned LM2BCPU+40 ptsall–8%614 ms
11Laya multilingualDecision encoder322MApple GPU+37 ptsall–17%15 ms
12cbjev multilingualDecision encoder322MApple GPU+37 ptsall (5 count)–15%20 ms
13Laya multilingual · fine-tunedDecision encoder322MApple GPU+36 ptsall–17%15 ms
14decider 0.8BFine-tuned LM0.8BCPU+36 ptsallEntailment16%279 ms
15Laya English · OllayaDecision encoder421MApple GPU+36 ptsall–18%32 ms
16Decision 1.0 EosFine-tuned LM0.8BCPU+36 ptsall–11%391 ms
17Laya multilingual · OllayaDecision encoder322MApple GPU+36 ptsall–18%13 ms
18Laya EnglishDecision encoder421MApple GPU+36 ptsall–20%35 ms
19jeff (GLiFormer)Zero-shot encoder435MApple GPU+35 ptsall–22%90 ms
20NLI zero-shot (DeBERTa)Zero-shot encoder435MCPU+35 ptsall–14%366 ms
21jevlike (trained per data set)Trained per data set307M + headApple GPU+30 ptsallPrompt injection27%14 ms
22Kev 0.8B · OllayaFine-tuned LM0.8BCPU+29 ptsall (6 count)–14%286 ms
23Von 1.1Zero-shot encoder395MCPU+24 ptsall–21%120 ms
24GLiClass instruct largeZero-shot encoder435MCPU+20 ptsall–24%123 ms
25Tacet SonataDecision encoder144MApple GPU+8 ptsall (9 count)–16%13 ms
26CLM-8BFrozen LM8BApple GPU+7 ptsall (1 can't answer)–34%114 ms
Yardsticks — simple rules, not models. Every model should beat them.
Word overlapyardstickYardstick—CPU+2 ptsall–6%0 ms
Most common answeryardstickYardstick—CPU−1 ptsall–1%0 ms

What stands out