systemone lab

How it works

The rules that keep the numbers honest.

The task
Each model is given a text, a question and a short list of possible answers, and must pick one with a confidence. It never writes free text. That is what makes models of very different sizes comparable.
Gain over guessing
Accuracy minus the score of always answering the most common label. A model at 75% looks fine until you learn that guessing scores 85%. Every score on this site sits beside its guessing score.
Right answers
Written by people, published with each data set, and never edited here. Data sets are sampled at a pinned revision so anyone can rebuild the same test.
Final exam
Each model is scored once on questions it was never tuned against. Practice runs are for building, and are never quoted.
95% range
Small tests are noisy. The range shows how far the true accuracy could plausibly be from the number (a Wilson interval on the questions scored).
Can't answer
Some models handle only yes/no and pick-one questions, not scales. A data set a model cannot answer at all counts as guessing (no gain) in its average, so a narrower model still ranks — against the others, not on a subset.
† Trained on the source
When a model's published training data includes the data set, its score there is in-distribution. It is shown and marked, and left out of that model's average.
Confidence off by
Every answer comes with a confidence. Group the answers by that confidence and compare each group's average confidence with how often it was actually right; the weighted average gap is the expected calibration error. 9% means a model that says 90% is right about 81% of the time. Lower is better: a model whose confidence can be trusted lets you act on it automatically above a threshold and send the rest to a person.
Head to head
On the same questions, count where only model A is right and where only model B is right. The exact McNemar test says how likely that split is if the two were really equal.
Speed
Median time per answered question, from runs made with nothing else on the machine (an Apple-silicon laptop with 64 GB). Hosted services include the network round trip.

Data sets

Data setTaskQuestionsGuessingSource and licence
News topic
Which newspaper section a story belongs in: world, sports, business, or science and technology.
pick one of 4
en
8025%fancyzhx/ag_news
revision eb185aade064
non-commercial research (AG's corpus)
Product review stars
How many stars, from 1 to 5, a reviewer gave. English, French, German and Spanish reviews in equal parts.
a scale of 5
en, fr, de, es (125 each)
8020%mteb/amazon_reviews_multi
revision c379a6705fec
Amazon Multilingual Reviews licence (research)
Yes/no from a passage
Answer a naturally occurring yes/no question using a Wikipedia passage — the model has to read the passage, not just the question.
yes / no
en
8050%google/boolq
revision 35b264d03638
CC BY-SA 3.0
Spam or not
Whether an email is unsolicited bulk mail. Real messages from a public corpus, some of them long.
yes / no
en
8055%SetFit/enron_spam
revision 1916f66c89d5
public corpus (Enron, FERC release)
Finance tweet stance
Whether a finance-related tweet is bullish, bearish or neutral.
pick one of 3
en
8034%zeroshot/twitter-financial-news-sentiment
revision ccbe24de388e
MIT
Financial news tone
The tone of a sentence from financial news — positive, negative or neutral — on sentences every annotator agreed on.
pick one of 3
en
8034%takala/financial_phrasebank
revision 8d3fe0c36d5f
CC BY-NC-SA 3.0
Intent routing
Which of many assistant scenarios a spoken-style request belongs to, in several languages. A wide choice, so guessing scores very low.
pick one of 16
en, fr, de, es, ru (100 each)
806%mteb/MassiveIntentClassification
revision 0f29dfaa4a2d
CC BY 4.0
Prompt injection
Whether a message tries to hijack an AI assistant's instructions.
yes / no
en, de
8050%deepset/prompt-injections
revision 4f61ecb038e9
Apache-2.0
Unfair contract terms
Whether a clause in a terms-of-service document is potentially unfair to the user.
yes / no
en
8055%coastalcph/lex_glue
revision c23fdff1a6bf
CC BY 4.0
Does one sentence follow from another?
Given two sentences, does the second follow from the first, contradict it, or neither? In several languages.
pick one of 3
en, fr, de, es, ru (100 each)
8038%facebook/xnli
revision b8dd5d7af511
CC BY-NC 4.0

Only scores are published. No text from the data sets appears on this site.