systemone lab

Models

What was compared. Every model is shown the same question, the same text and the same list of answers, and must pick one. The row of squares under each model is its result on every data set, in the leaderboard's order: blue above guessing, red below, † where it trained on that data set's source.

Decision encoder

Laya multilingual

322MApple GPU

Convai's multilingual Laya: an mmBERT-base encoder with a head that scores each option, covering 100+ languages. Used exactly as released, without calibration. source ↗

average gain +37 pts · confidence off by 17% · 15 ms per answer

cbjev multilingual

322MApple GPU

A community fine-tune of Laya multilingual on a 26-language entailment mix plus reviews, tweets, topics and toxicity. Run through Laya's own sequence builder. source ↗

†††††

average gain +37 pts over 5 data sets · confidence off by 15% · 20 ms per answer

Laya multilingual · fine-tuned

322MApple GPU

Laya multilingual after a small supervised fine-tune of its head on a private labelled set unrelated to these data sets. Here to show whether tuning for one job helps or hurts everywhere else. source ↗

average gain +36 pts · confidence off by 17% · 15 ms per answer

Laya English · Ollaya

421MApple GPU

The same English Laya weights, served by Ollaya's runtime on Apple Metal. source ↗

average gain +36 pts · confidence off by 18% · 32 ms per answer

Laya multilingual · Ollaya

322MApple GPU

The same multilingual Laya weights, served by Ollaya's runtime on Apple Metal instead of in-process PyTorch. Any difference from the row above is runtime, window length and speed. source ↗

average gain +36 pts · confidence off by 18% · 13 ms per answer

Laya English

421MApple GPU

Convai's English Laya on a ModernBERT-large encoder. Used exactly as released. source ↗

average gain +36 pts · confidence off by 20% · 35 ms per answer

Tacet Sonata

144MApple GPU

A 144M typed-decision encoder: mmBERT-small, a two-layer transformer head and Laya's option scorer. All of a text's questions are packed into one sequence. source ↗

†

average gain +8 pts over 9 data sets · confidence off by 16% · 13 ms per answer

Zero-shot encoder

jeff (GLiFormer)

435MApple GPU

A DeBERTa-v3-large GLiFormer zero-shot classifier: an independent score per label, renormalised over the options. source ↗

average gain +35 pts · confidence off by 22% · 90 ms per answer

NLI zero-shot (DeBERTa)

435MCPU

A DeBERTa-v3-large entailment model used the classic zero-shot way: every option becomes a hypothesis, scored for entailment against the text. source ↗

average gain +35 pts · confidence off by 14% · 366 ms per answer

Von 1.1

395MCPU

A ModernBERT-large model that reads the instructions, the text and every option after its own marker, and scores each option at its marker. source ↗

average gain +24 pts · confidence off by 21% · 120 ms per answer

GLiClass instruct large

435MCPU

Knowledgator's instruction-following zero-shot classifier on DeBERTa-v3-large. Every option is a label span pooled and scored against the text. source ↗

average gain +20 pts · confidence off by 24% · 123 ms per answer

Fine-tuned LM

Winnow E4B

E4BApple GPU

A Gemma 4 E4B fine-tune. Options become letters and the answer is the next-token probability of each letter. Run as the author's Q8 GGUF on llama.cpp. source ↗

average gain +47 pts · confidence off by 10% · 184 ms per answer · best on 1 data set

JevK5 4B

4BApple GPU

A Qwen3.5-4B fine-tune that reads text, question and options as one JSON message and answers with a letter; the probability of each letter is the answer. Run as the author's Q8 GGUF on llama.cpp. source ↗

average gain +44 pts · confidence off by 10% · 181 ms per answer

Jev-Style 0.8B

0.8BApple GPU

A fully fine-tuned Qwen3.5-0.8B that reads text, question and options once and scores each option at its own slot. Run on MLX. source ↗

††

average gain +42 pts over 8 data sets · confidence off by 12% · 39 ms per answer · best on 3 data sets

Kev 9B · Ollaya

9BCPU

The largest Kev, on Qwen3.5-9B, served by Ollaya on the CPU. source ↗

††††

average gain +42 pts over 6 data sets · confidence off by 14% · 3.2 s per answer

Kev 4B · MLX

4BApple GPU

Jared Palmer's Kev: a Qwen3.5-4B base with a rank-16 LoRA merged in and a pointer head that scores each option at its own span. Run by Kev's own code on MLX. source ↗

††††

average gain +41 pts over 6 data sets · confidence off by 14% · 131 ms per answer

Kev 4B · Ollaya

4BCPU

Kev on Qwen3.5-4B, served by Ollaya on the CPU. Same design as the MLX row; a different runtime. source ↗

††††

average gain +41 pts over 6 data sets · confidence off by 14% · 1.8 s per answer

decider 2B

2BCPU

Mapika's decider on a Qwen3.5-2B base: text, question and lettered options in, the option letters' next-token logits out. source ↗

average gain +40 pts · confidence off by 8% · 614 ms per answer

decider 0.8B

0.8BCPU

Mapika's smallest decider: Qwen3.5-0.8B reading the logits of the option letters at an answer slot. One forward pass per question, no generation. source ↗

average gain +36 pts · confidence off by 16% · 279 ms per answer · best on 1 data set

Decision 1.0 Eos

0.8BCPU

The vLLM Semantic Router contributors' fully fine-tuned Qwen3.5-0.8B with a bilinear head that scores each option's last token against the end of the question. source ↗

average gain +36 pts · confidence off by 11% · 391 ms per answer

Kev 0.8B · Ollaya

0.8BCPU

The smallest Kev: LoRA plus pointer head on Qwen3.5-0.8B, served by Ollaya on the CPU. source ↗

††††

average gain +29 pts over 6 data sets · confidence off by 14% · 286 ms per answer

Frozen LM

ruling 35B-A3B

35B (3B active)Apple GPU

Qwen3.6-35B-A3B, a frozen mixture-of-experts chat model with 3B active parameters. It reads the text once, answers each question from the letter logits, and averages three orderings of the options so position cannot bias it. source ↗

average gain +47 pts · confidence off by 12% · 391 ms per answer · best on 1 data set

SemIf 4B

4BApple GPU

Qwen3.5-4B, a general chat model used as is: it sees the text and the options as letters, and the softmax over the letter logits is its answer. One forward pass, nothing generated. source ↗

average gain +43 pts · confidence off by 12% · 171 ms per answer

CLM-8B

8BApple GPU

A bi-encoder on a frozen Qwen3-8B: the text-plus-question and each option are embedded separately, two small trained heads project both, and a softmax over cosines is the answer. source ↗

average gain +7 pts · confidence off by 34% · 114 ms per answer

Trained per data set

jevlike (trained per data set)

307M + headApple GPU

No checkpoint ships. Before each data set, a fresh attention head is trained on that set's training rows over a frozen multilingual encoder. Unlike every other row, it has seen labelled examples of the task — from the training split, never the scored one. source ↗

average gain +30 pts · confidence off by 27% · 14 ms per answer · best on 1 data set

Hosted

Jev (hosted)

undisclosedhosted service

TypeSafe's closed model, reached over the network through OpenRouter. The reference every open alternative here is measured against, and the only row that does not run locally. source ↗

average gain +49 pts · confidence off by 9% · 101 ms per answer · best on 3 data sets

Yardstick

Word overlap

—CPUyardstick

Not a model: picks the answer whose description shares the most words with the text. A model that cannot beat it is matching keywords, not reading.

average gain +2 pts · confidence off by 6% · 0 ms per answer

Most common answer

—CPUyardstick

Not a model: answers with the training rows' label frequencies and never reads the text. Anything that cannot beat it is not reading either.

average gain −1 pts · confidence off by 1% · 0 ms per answer