What was compared. Every model is shown the same question, the same text and the same list of answers, and
must pick one. The row of squares under each model is its result on every data set, in the leaderboard's
order: blue above guessing, red below, † where it trained on that data set's source.
Laya multilingual
322MApple GPU
Convai's multilingual Laya: an mmBERT-base encoder with a head that scores each option, covering 100+ languages. Used exactly as released, without calibration. source ↗
average gain +37 pts · confidence off by 17% · 15 ms per answer
cbjev multilingual
322MApple GPU
A community fine-tune of Laya multilingual on a 26-language entailment mix plus reviews, tweets, topics and toxicity. Run through Laya's own sequence builder. source ↗
†††††
average gain +37 pts over 5 data sets · confidence off by 15% · 20 ms per answer
Laya multilingual · fine-tuned
322MApple GPU
Laya multilingual after a small supervised fine-tune of its head on a private labelled set unrelated to these data sets. Here to show whether tuning for one job helps or hurts everywhere else. source ↗
average gain +36 pts · confidence off by 17% · 15 ms per answer
Laya English · Ollaya
421MApple GPU
The same English Laya weights, served by Ollaya's runtime on Apple Metal. source ↗
average gain +36 pts · confidence off by 18% · 32 ms per answer
Laya multilingual · Ollaya
322MApple GPU
The same multilingual Laya weights, served by Ollaya's runtime on Apple Metal instead of in-process PyTorch. Any difference from the row above is runtime, window length and speed. source ↗
average gain +36 pts · confidence off by 18% · 13 ms per answer
Laya English
421MApple GPU
Convai's English Laya on a ModernBERT-large encoder. Used exactly as released. source ↗
average gain +36 pts · confidence off by 20% · 35 ms per answer
Tacet Sonata
144MApple GPU
A 144M typed-decision encoder: mmBERT-small, a two-layer transformer head and Laya's option scorer. All of a text's questions are packed into one sequence. source ↗
†
average gain +8 pts over 9 data sets · confidence off by 16% · 13 ms per answer
jeff (GLiFormer)
435MApple GPU
A DeBERTa-v3-large GLiFormer zero-shot classifier: an independent score per label, renormalised over the options. source ↗
average gain +35 pts · confidence off by 22% · 90 ms per answer
NLI zero-shot (DeBERTa)
435MCPU
A DeBERTa-v3-large entailment model used the classic zero-shot way: every option becomes a hypothesis, scored for entailment against the text. source ↗
average gain +35 pts · confidence off by 14% · 366 ms per answer
Von 1.1
395MCPU
A ModernBERT-large model that reads the instructions, the text and every option after its own marker, and scores each option at its marker. source ↗
average gain +24 pts · confidence off by 21% · 120 ms per answer
GLiClass instruct large
435MCPU
Knowledgator's instruction-following zero-shot classifier on DeBERTa-v3-large. Every option is a label span pooled and scored against the text. source ↗
average gain +20 pts · confidence off by 24% · 123 ms per answer
Winnow E4B
E4BApple GPU
A Gemma 4 E4B fine-tune. Options become letters and the answer is the next-token probability of each letter. Run as the author's Q8 GGUF on llama.cpp. source ↗
average gain +47 pts · confidence off by 10% · 184 ms per answer · best on 1 data set
JevK5 4B
4BApple GPU
A Qwen3.5-4B fine-tune that reads text, question and options as one JSON message and answers with a letter; the probability of each letter is the answer. Run as the author's Q8 GGUF on llama.cpp. source ↗
average gain +44 pts · confidence off by 10% · 181 ms per answer
Jev-Style 0.8B
0.8BApple GPU
A fully fine-tuned Qwen3.5-0.8B that reads text, question and options once and scores each option at its own slot. Run on MLX. source ↗
††
average gain +42 pts over 8 data sets · confidence off by 12% · 39 ms per answer · best on 3 data sets
Kev 9B · Ollaya
9BCPU
The largest Kev, on Qwen3.5-9B, served by Ollaya on the CPU. source ↗
††††
average gain +42 pts over 6 data sets · confidence off by 14% · 3.2 s per answer
Kev 4B · MLX
4BApple GPU
Jared Palmer's Kev: a Qwen3.5-4B base with a rank-16 LoRA merged in and a pointer head that scores each option at its own span. Run by Kev's own code on MLX. source ↗
††††
average gain +41 pts over 6 data sets · confidence off by 14% · 131 ms per answer
Kev 4B · Ollaya
4BCPU
Kev on Qwen3.5-4B, served by Ollaya on the CPU. Same design as the MLX row; a different runtime. source ↗
††††
average gain +41 pts over 6 data sets · confidence off by 14% · 1.8 s per answer
decider 2B
2BCPU
Mapika's decider on a Qwen3.5-2B base: text, question and lettered options in, the option letters' next-token logits out. source ↗
average gain +40 pts · confidence off by 8% · 614 ms per answer
decider 0.8B
0.8BCPU
Mapika's smallest decider: Qwen3.5-0.8B reading the logits of the option letters at an answer slot. One forward pass per question, no generation. source ↗
average gain +36 pts · confidence off by 16% · 279 ms per answer · best on 1 data set
Decision 1.0 Eos
0.8BCPU
The vLLM Semantic Router contributors' fully fine-tuned Qwen3.5-0.8B with a bilinear head that scores each option's last token against the end of the question. source ↗
average gain +36 pts · confidence off by 11% · 391 ms per answer
Kev 0.8B · Ollaya
0.8BCPU
The smallest Kev: LoRA plus pointer head on Qwen3.5-0.8B, served by Ollaya on the CPU. source ↗
††††
average gain +29 pts over 6 data sets · confidence off by 14% · 286 ms per answer
ruling 35B-A3B
35B (3B active)Apple GPU
Qwen3.6-35B-A3B, a frozen mixture-of-experts chat model with 3B active parameters. It reads the text once, answers each question from the letter logits, and averages three orderings of the options so position cannot bias it. source ↗
average gain +47 pts · confidence off by 12% · 391 ms per answer · best on 1 data set
SemIf 4B
4BApple GPU
Qwen3.5-4B, a general chat model used as is: it sees the text and the options as letters, and the softmax over the letter logits is its answer. One forward pass, nothing generated. source ↗
average gain +43 pts · confidence off by 12% · 171 ms per answer
CLM-8B
8BApple GPU
A bi-encoder on a frozen Qwen3-8B: the text-plus-question and each option are embedded separately, two small trained heads project both, and a softmax over cosines is the answer. source ↗
average gain +7 pts · confidence off by 34% · 114 ms per answer
jevlike (trained per data set)
307M + headApple GPU
No checkpoint ships. Before each data set, a fresh attention head is trained on that set's training rows over a frozen multilingual encoder. Unlike every other row, it has seen labelled examples of the task — from the training split, never the scored one. source ↗
average gain +30 pts · confidence off by 27% · 14 ms per answer · best on 1 data set
Jev (hosted)
undisclosedhosted service
TypeSafe's closed model, reached over the network through OpenRouter. The reference every open alternative here is measured against, and the only row that does not run locally. source ↗
average gain +49 pts · confidence off by 9% · 101 ms per answer · best on 3 data sets
Word overlap
—CPUyardstick
Not a model: picks the answer whose description shares the most words with the text. A model that cannot beat it is matching keywords, not reading.
average gain +2 pts · confidence off by 6% · 0 ms per answer
Most common answer
—CPUyardstick
Not a model: answers with the training rows' label frequencies and never reads the text. Anything that cannot beat it is not reading either.
average gain −1 pts · confidence off by 1% · 0 ms per answer