How it works
The rules that keep the numbers honest.
- The task
- Each model is given a text, a question and a short list of possible answers, and must pick one with a confidence. It never writes free text. That is what makes models of very different sizes comparable.
- Gain over guessing
- Accuracy minus the score of always answering the most common label. A model at 75% looks fine until you learn that guessing scores 85%. Every score on this site sits beside its guessing score.
- Right answers
- Written by people, published with each data set, and never edited here. Data sets are sampled at a pinned revision so anyone can rebuild the same test.
- Final exam
- Each model is scored once on questions it was never tuned against. Practice runs are for building, and are never quoted.
- 95% range
- Small tests are noisy. The range shows how far the true accuracy could plausibly be from the number (a Wilson interval on the questions scored).
- Can't answer
- Some models handle only yes/no and pick-one questions, not scales. A data set a model cannot answer at all counts as guessing (no gain) in its average, so a narrower model still ranks — against the others, not on a subset.
- † Trained on the source
- When a model's published training data includes the data set, its score there is in-distribution. It is shown and marked, and left out of that model's average.
- Confidence off by
- Every answer comes with a confidence. Group the answers by that confidence and compare each group's average confidence with how often it was actually right; the weighted average gap is the expected calibration error. 9% means a model that says 90% is right about 81% of the time. Lower is better: a model whose confidence can be trusted lets you act on it automatically above a threshold and send the rest to a person.
- Head to head
- On the same questions, count where only model A is right and where only model B is right. The exact McNemar test says how likely that split is if the two were really equal.
- Speed
- Median time per answered question, from runs made with nothing else on the machine (an Apple-silicon laptop with 64 GB). Hosted services include the network round trip.
Data sets
| Data set | Task | Questions | Guessing | Source and licence |
|---|---|---|---|---|
| News topic Which newspaper section a story belongs in: world, sports, business, or science and technology. | pick one of 4 en | 80 | 25% | fancyzhx/ag_news revision eb185aade064 non-commercial research (AG's corpus) |
| Product review stars How many stars, from 1 to 5, a reviewer gave. English, French, German and Spanish reviews in equal parts. | a scale of 5 en, fr, de, es (125 each) | 80 | 20% | mteb/amazon_reviews_multi revision c379a6705fec Amazon Multilingual Reviews licence (research) |
| Yes/no from a passage Answer a naturally occurring yes/no question using a Wikipedia passage — the model has to read the passage, not just the question. | yes / no en | 80 | 50% | google/boolq revision 35b264d03638 CC BY-SA 3.0 |
| Spam or not Whether an email is unsolicited bulk mail. Real messages from a public corpus, some of them long. | yes / no en | 80 | 55% | SetFit/enron_spam revision 1916f66c89d5 public corpus (Enron, FERC release) |
| Finance tweet stance Whether a finance-related tweet is bullish, bearish or neutral. | pick one of 3 en | 80 | 34% | zeroshot/twitter-financial-news-sentiment revision ccbe24de388e MIT |
| Financial news tone The tone of a sentence from financial news — positive, negative or neutral — on sentences every annotator agreed on. | pick one of 3 en | 80 | 34% | takala/financial_phrasebank revision 8d3fe0c36d5f CC BY-NC-SA 3.0 |
| Intent routing Which of many assistant scenarios a spoken-style request belongs to, in several languages. A wide choice, so guessing scores very low. | pick one of 16 en, fr, de, es, ru (100 each) | 80 | 6% | mteb/MassiveIntentClassification revision 0f29dfaa4a2d CC BY 4.0 |
| Prompt injection Whether a message tries to hijack an AI assistant's instructions. | yes / no en, de | 80 | 50% | deepset/prompt-injections revision 4f61ecb038e9 Apache-2.0 |
| Unfair contract terms Whether a clause in a terms-of-service document is potentially unfair to the user. | yes / no en | 80 | 55% | coastalcph/lex_glue revision c23fdff1a6bf CC BY 4.0 |
| Does one sentence follow from another? Given two sentences, does the second follow from the first, contradict it, or neither? In several languages. | pick one of 3 en, fr, de, es, ru (100 each) | 80 | 38% | facebook/xnli revision b8dd5d7af511 CC BY-NC 4.0 |
Only scores are published. No text from the data sets appears on this site.