systemone lab

By data set

One task at a time: every model's accuracy with its 95% range, against the line for guessing the most common answer. Pick a data set.

News topic

Which newspaper section a story belongs in: world, sports, business, or science and technology.

pick one of 480 questionsguessing scores 25%en

Jev-Style 0.8B is the most accurate at 98%; 21 of 21 models beat guessing.

guessing (25%)95% rangemodelyardsticktrained on this data set's source (†)

Source: fancyzhx/ag_news · non-commercial research (AG's corpus)

Product review stars

How many stars, from 1 to 5, a reviewer gave. English, French, German and Spanish reviews in equal parts.

a scale of 580 questionsguessing scores 20%en, fr, de, es (125 each)

Jev (hosted) is the most accurate at 62%; 20 of 20 models beat guessing.

guessing (20%)95% rangemodelyardsticktrained on this data set's source (†)

Source: mteb/amazon_reviews_multi · Amazon Multilingual Reviews licence (research)

Yes/no from a passage

Answer a naturally occurring yes/no question using a Wikipedia passage — the model has to read the passage, not just the question.

yes / no80 questionsguessing scores 50%en

Jev (hosted) is the most accurate at 95%; 21 of 21 models beat guessing.

guessing (50%)95% rangemodelyardsticktrained on this data set's source (†)

Source: google/boolq · CC BY-SA 3.0

Spam or not

Whether an email is unsolicited bulk mail. Real messages from a public corpus, some of them long.

yes / no80 questionsguessing scores 55%en

Jev-Style 0.8B is the most accurate at 100%; 23 of 25 models beat guessing.

guessing (55%)95% rangemodelyardsticktrained on this data set's source (†)

Source: SetFit/enron_spam · public corpus (Enron, FERC release)

Finance tweet stance

Whether a finance-related tweet is bullish, bearish or neutral.

pick one of 380 questionsguessing scores 34%en

ruling 35B-A3B is the most accurate at 80%; 25 of 26 models beat guessing.

guessing (34%)95% rangemodelyardsticktrained on this data set's source (†)

Source: zeroshot/twitter-financial-news-sentiment · MIT

Financial news tone

The tone of a sentence from financial news — positive, negative or neutral — on sentences every annotator agreed on.

pick one of 380 questionsguessing scores 34%en

Winnow E4B is the most accurate at 99%; 26 of 26 models beat guessing.

guessing (34%)95% rangemodelyardsticktrained on this data set's source (†)

Source: takala/financial_phrasebank · CC BY-NC-SA 3.0

Intent routing

Which of many assistant scenarios a spoken-style request belongs to, in several languages. A wide choice, so guessing scores very low.

pick one of 1680 questionsguessing scores 6%en, fr, de, es, ru (100 each)

Jev (hosted) is the most accurate at 96%; 24 of 24 models beat guessing.

guessing (6%)95% rangemodelyardsticktrained on this data set's source (†)

Source: mteb/MassiveIntentClassification · CC BY 4.0

Prompt injection

Whether a message tries to hijack an AI assistant's instructions.

yes / no80 questionsguessing scores 50%en, de

jevlike (trained per data set) is the most accurate at 95%; 25 of 26 models beat guessing.

guessing (50%)95% rangemodelyardsticktrained on this data set's source (†)

Source: deepset/prompt-injections · Apache-2.0

Unfair contract terms

Whether a clause in a terms-of-service document is potentially unfair to the user.

yes / no80 questionsguessing scores 55%en

Jev-Style 0.8B is the most accurate at 92%; 18 of 26 models beat guessing.

guessing (55%)95% rangemodelyardsticktrained on this data set's source (†)

Source: coastalcph/lex_glue · CC BY 4.0

Does one sentence follow from another?

Given two sentences, does the second follow from the first, contradict it, or neither? In several languages.

pick one of 380 questionsguessing scores 38%en, fr, de, es, ru (100 each)

decider 0.8B is the most accurate at 81%; 14 of 20 models beat guessing.

guessing (38%)95% rangemodelyardsticktrained on this data set's source (†)

Source: facebook/xnli · CC BY-NC 4.0

All data sets at once

Accuracy, shaded by how far above (blue) or below (red) guessing it is. Click a heading to sort.

ModelNews topicReview starsYes/no readingSpamFinance tweetsFinancial newsIntent routingPrompt injectionUnfair termsEntailmentAvg gain
Guessing the most common answer25%20%50%55%34%34%6%50%55%38%0 pts
Jev (hosted)89%62%95%99%74%98%96%80%84%78%+49 pts
ruling 35B-A3B88%57%91%99%80%90%90%89%74%78%+47 pts
Winnow E4B85%61%86%99%79%99%91%86%78%70%+47 pts
JevK5 4B88%61%88%95%69%95%92%80%64%75%+44 pts
SemIf 4B90%55%85%92%75%95%84%72%72%74%+43 pts
Jev-Style 0.8B98%50%82%100%76%84%90%†75%92%88%†+42 pts
Kev 9B · Ollaya95%†76%†92%†91%76%95%89%68%65%79%†+42 pts
Kev 4B · MLX95%†64%†90%†96%60%86%89%70%76%80%†+41 pts
Kev 4B · Ollaya95%†62%†90%†96%60%86%89%70%76%80%†+41 pts
decider 2B94%52%84%100%72%72%86%69%61%80%+40 pts
Laya multilingual96%40%74%98%69%94%70%69%54%75%+37 pts
cbjev multilingual95%†48%†79%†97%†69%92%70%76%56%79%†+37 pts
Laya multilingual · fine-tuned96%41%74%96%68%89%69%69%55%74%+36 pts
decider 0.8B91%59%84%99%59%35%81%65%74%81%+36 pts
Laya English · Ollaya92%36%85%100%60%98%52%81%56%66%+36 pts
Decision 1.0 Eos90%56%71%78%74%82%79%78%51%65%+36 pts
Laya multilingual · Ollaya96%30%74%97%69%94%70%69%54%75%+36 pts
Laya English92%34%84%99%60%98%52%81%56%66%+36 pts
jeff (GLiFormer)82%62%70%95%74%94%68%59%55%59%+35 pts
NLI zero-shot (DeBERTa)91%54%90%87%70%85%51%84%72%31%+35 pts
jevlike (trained per data set)89%21%54%90%75%81%51%95%79%29%+30 pts
Kev 0.8B · Ollaya94%†64%†85%†69%64%72%78%54%72%74%†+29 pts
Von 1.185%24%74%76%45%91%30%74%74%35%+24 pts
GLiClass instruct large88%41%65%65%51%76%41%57%45%36%+20 pts
Tacet Sonata74%32%59%54%40%49%50%†50%44%31%+8 pts
CLM-8B55%can't55%44%32%46%31%84%45%29%+7 pts
Word overlap36%19%50%59%35%34%11%50%60%38%+2 pts
Most common answer25%20%50%55%32%34%6%50%55%31%−1 pts

† trained on this data set's source. Italic rows are yardsticks.