Preliminary numbers. These come from the practice split. The final exam — one attempt per model on questions never tuned against — has not been run yet, so rankings may still move.
Head to head
Two accuracies can differ by ten points and still mean nothing on a small test. What separates two models is
how many questions one gets right that the other gets wrong, and whether that split could be chance.
against
p is the exact McNemar test on the questions where exactly one of the two is right.
Below 0.05 we call the gap real.