By data set
One task at a time: every model's accuracy with its 95% range, against the line for guessing the most common answer. Pick a data set.
News topic
Which newspaper section a story belongs in: world, sports, business, or science and technology.
Jev-Style 0.8B is the most accurate at 98%; 21 of 21 models beat guessing.
Source: fancyzhx/ag_news · non-commercial research (AG's corpus)
Product review stars
How many stars, from 1 to 5, a reviewer gave. English, French, German and Spanish reviews in equal parts.
Jev (hosted) is the most accurate at 62%; 20 of 20 models beat guessing.
Source: mteb/amazon_reviews_multi · Amazon Multilingual Reviews licence (research)
Yes/no from a passage
Answer a naturally occurring yes/no question using a Wikipedia passage — the model has to read the passage, not just the question.
Jev (hosted) is the most accurate at 95%; 21 of 21 models beat guessing.
Source: google/boolq · CC BY-SA 3.0
Spam or not
Whether an email is unsolicited bulk mail. Real messages from a public corpus, some of them long.
Jev-Style 0.8B is the most accurate at 100%; 23 of 25 models beat guessing.
Source: SetFit/enron_spam · public corpus (Enron, FERC release)
Finance tweet stance
Whether a finance-related tweet is bullish, bearish or neutral.
ruling 35B-A3B is the most accurate at 80%; 25 of 26 models beat guessing.
Source: zeroshot/twitter-financial-news-sentiment · MIT
Financial news tone
The tone of a sentence from financial news — positive, negative or neutral — on sentences every annotator agreed on.
Winnow E4B is the most accurate at 99%; 26 of 26 models beat guessing.
Source: takala/financial_phrasebank · CC BY-NC-SA 3.0
Intent routing
Which of many assistant scenarios a spoken-style request belongs to, in several languages. A wide choice, so guessing scores very low.
Jev (hosted) is the most accurate at 96%; 24 of 24 models beat guessing.
Source: mteb/MassiveIntentClassification · CC BY 4.0
Prompt injection
Whether a message tries to hijack an AI assistant's instructions.
jevlike (trained per data set) is the most accurate at 95%; 25 of 26 models beat guessing.
Source: deepset/prompt-injections · Apache-2.0
Unfair contract terms
Whether a clause in a terms-of-service document is potentially unfair to the user.
Jev-Style 0.8B is the most accurate at 92%; 18 of 26 models beat guessing.
Source: coastalcph/lex_glue · CC BY 4.0
Does one sentence follow from another?
Given two sentences, does the second follow from the first, contradict it, or neither? In several languages.
decider 0.8B is the most accurate at 81%; 14 of 20 models beat guessing.
Source: facebook/xnli · CC BY-NC 4.0
All data sets at once
Accuracy, shaded by how far above (blue) or below (red) guessing it is. Click a heading to sort.
| Model | News topic | Review stars | Yes/no reading | Spam | Finance tweets | Financial news | Intent routing | Prompt injection | Unfair terms | Entailment | Avg gain |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Guessing the most common answer | 25% | 20% | 50% | 55% | 34% | 34% | 6% | 50% | 55% | 38% | 0 pts |
| Jev (hosted) | 89% | 62% | 95% | 99% | 74% | 98% | 96% | 80% | 84% | 78% | +49 pts |
| ruling 35B-A3B | 88% | 57% | 91% | 99% | 80% | 90% | 90% | 89% | 74% | 78% | +47 pts |
| Winnow E4B | 85% | 61% | 86% | 99% | 79% | 99% | 91% | 86% | 78% | 70% | +47 pts |
| JevK5 4B | 88% | 61% | 88% | 95% | 69% | 95% | 92% | 80% | 64% | 75% | +44 pts |
| SemIf 4B | 90% | 55% | 85% | 92% | 75% | 95% | 84% | 72% | 72% | 74% | +43 pts |
| Jev-Style 0.8B | 98% | 50% | 82% | 100% | 76% | 84% | 90%† | 75% | 92% | 88%† | +42 pts |
| Kev 9B · Ollaya | 95%† | 76%† | 92%† | 91% | 76% | 95% | 89% | 68% | 65% | 79%† | +42 pts |
| Kev 4B · MLX | 95%† | 64%† | 90%† | 96% | 60% | 86% | 89% | 70% | 76% | 80%† | +41 pts |
| Kev 4B · Ollaya | 95%† | 62%† | 90%† | 96% | 60% | 86% | 89% | 70% | 76% | 80%† | +41 pts |
| decider 2B | 94% | 52% | 84% | 100% | 72% | 72% | 86% | 69% | 61% | 80% | +40 pts |
| Laya multilingual | 96% | 40% | 74% | 98% | 69% | 94% | 70% | 69% | 54% | 75% | +37 pts |
| cbjev multilingual | 95%† | 48%† | 79%† | 97%† | 69% | 92% | 70% | 76% | 56% | 79%† | +37 pts |
| Laya multilingual · fine-tuned | 96% | 41% | 74% | 96% | 68% | 89% | 69% | 69% | 55% | 74% | +36 pts |
| decider 0.8B | 91% | 59% | 84% | 99% | 59% | 35% | 81% | 65% | 74% | 81% | +36 pts |
| Laya English · Ollaya | 92% | 36% | 85% | 100% | 60% | 98% | 52% | 81% | 56% | 66% | +36 pts |
| Decision 1.0 Eos | 90% | 56% | 71% | 78% | 74% | 82% | 79% | 78% | 51% | 65% | +36 pts |
| Laya multilingual · Ollaya | 96% | 30% | 74% | 97% | 69% | 94% | 70% | 69% | 54% | 75% | +36 pts |
| Laya English | 92% | 34% | 84% | 99% | 60% | 98% | 52% | 81% | 56% | 66% | +36 pts |
| jeff (GLiFormer) | 82% | 62% | 70% | 95% | 74% | 94% | 68% | 59% | 55% | 59% | +35 pts |
| NLI zero-shot (DeBERTa) | 91% | 54% | 90% | 87% | 70% | 85% | 51% | 84% | 72% | 31% | +35 pts |
| jevlike (trained per data set) | 89% | 21% | 54% | 90% | 75% | 81% | 51% | 95% | 79% | 29% | +30 pts |
| Kev 0.8B · Ollaya | 94%† | 64%† | 85%† | 69% | 64% | 72% | 78% | 54% | 72% | 74%† | +29 pts |
| Von 1.1 | 85% | 24% | 74% | 76% | 45% | 91% | 30% | 74% | 74% | 35% | +24 pts |
| GLiClass instruct large | 88% | 41% | 65% | 65% | 51% | 76% | 41% | 57% | 45% | 36% | +20 pts |
| Tacet Sonata | 74% | 32% | 59% | 54% | 40% | 49% | 50%† | 50% | 44% | 31% | +8 pts |
| CLM-8B | 55% | can't | 55% | 44% | 32% | 46% | 31% | 84% | 45% | 29% | +7 pts |
| Word overlap | 36% | 19% | 50% | 59% | 35% | 34% | 11% | 50% | 60% | 38% | +2 pts |
| Most common answer | 25% | 20% | 50% | 55% | 32% | 34% | 6% | 50% | 55% | 31% | −1 pts |
† trained on this data set's source. Italic rows are yardsticks.