What else we measure
Datum ingests every benchmark series in Epoch's archive. The Board publishes 14. These are the 66 it does not, and why.
A score with no evaluation date could have been run at any point, including after the questions leaked into a model's training data — and the gap between release and evaluation is exactly where contamination lives. A score with no error bar cannot tell you whether the distance between two models is real or noise. Series missing either are held, not published.
They still earn their place. If a model leads a Board house and tops a stack of series measured independently, that agreement is evidence. If it leads on the Board and nowhere else, that is worth knowing too. So this page counts agreement — and never turns it into a rank, because none of these series is measured well enough to support one.
17 of 66
of these series are topped by a model that also leads a house on the Board. That is a count of agreement, not a score — it has no threshold and no pass mark, because a threshold would be a verdict these numbers cannot carry.
Every series we hold
| Series | Coverage | Top scorer | Best score | Why it is not on the Board |
|---|---|---|---|---|
| dtbenchsaturating | 160 models · 161 scores | claude-fable-5 | 98.4% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| mmlu | 136 models · 136 scores | gpt-4o-2024-11-20 | 88.1% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| weirdml | 135 models · 169 scores | claude-fable-5-1✓ | 92.9% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| critpt | 123 models · 177 scores | gpt-5.6-sol✓ | 32.3% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| lmca | 122 models · 123 scores | claude-opus-5 | 63.3 | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| scicode | 113 models · 165 scores | claude-fable-5-1✓ | 62.0% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| ale bench | 102 models · 112 scores | gpt-5.6-sol✓ | 2177 | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| webdev arenaarena | 94 models · 120 scores | gpt-6-astra✓ | 1797 | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| gsm8k | 93 models · 93 scores | DeepSeek-Coder-V2-Instruct | 94.5% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| simplebench | 87 models · 102 scores | claude-fable-5 | 81.9% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| arc agisaturating | 83 models · 213 scores | gpt-6-astra✓ | 98.5% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| arc agi 2 | 80 models · 203 scores | gpt-6-astra✓ | 95.0% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| wino grande | 80 models · 80 scores | Llama-3.1-405B | 89.2% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| forecastbench | 78 models · 81 scores | o3-2025-04-16 | 62.5 | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| arc ai2 | 77 models · 77 scores | Llama-3.1-405B | 95.3% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| bool q | 77 models · 77 scores | T5-11B | 91.2% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| hella swag | 76 models · 76 scores | gpt-4-32k-0314 | 95.3% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| piqa | 60 models · 60 scores | gpt-4o-mini-2024-07-18 | 88.7% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| aider polyglot | 58 models · 71 scores | gpt-5-2025-08-07✓ | 88 | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| vending bench 2 | 55 models · 60 scores | claude-opus-5 | 11182 | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| apex agents | 54 models · 65 scores | gpt-6-astra✓ | 62.4% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| bbh | 50 models · 50 scores | gemini-1.5-pro-001 | 89.2% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| live bench | 50 models · 52 scores | gemini-2.5-pro-exp-03-25 | 82.35 | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| video mme | 50 models · 50 scores | video-SALMONN-2plus | 79.7% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| terminalbench | 45 models · 146 scores | gpt-5.5 | 84.7% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| lech mazur writing | 44 models · 49 scores | gpt-5-2025-08-07✓ | 8.6 | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| metr time horizons | 43 models · 50 scores | claude-mythos-preview-early | 1045 | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| hle | 42 models · 46 scores | claude-fable-5-1✓ | 46.5% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| open book qa | 42 models · 42 scores | Phi-3-mini-4k-instruct | 88.0% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| enigma eval | 41 models · 46 scores | claude-fable-5 | 39.3% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| trivia qa | 39 models · 39 scores | Llama-2-70b-hf | 87.6% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| balrog | 32 models · 35 scores | gemini-3-pro-preview | 58.1% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| frontiercode | 30 models · 30 scores | claude-fable-5 | 53.5% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| vpct | 28 models · 38 scores | gemini-3-pro-preview | 91.0% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| geobench | 27 models · 32 scores | gemini-3-flash-preview | 4333 | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| gso | 27 models · 27 scores | claude-opus-4-8 | 47.1% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| deepswe | 26 models · 26 scores | gpt-6-astra✓ | 74.1% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| lambada | 26 models · 26 scores | Megatron-Turing NLG 530B | 87.2% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| deepresearchbench | 25 models · 41 scores | claude-opus-4-6 | 55.3% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| gdp pdf | 25 models · 39 scores | gpt-5.6-sol✓ | 30.7% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| blueprint bench 2 | 24 models · 24 scores | claude-fable-5-1✓ | 41.9% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| gbaeval | 23 models · 23 scores | claude-opus-5 | 79.6% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| science qasaturating | 23 models · 23 scores | Phi-3.5-vision-instruct | 91.3% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| surface evolver bench | 23 models · 26 scores | kimi-k3 | 95.0% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| cursorbench | 22 models · 74 scores | claude-fable-5-1✓ | 73.4% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| cl bench | 21 models · 23 scores | gpt-5.4-2026-03-05 | 27.9% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| cybench | 21 models · 22 scores | claude-opus-4-6 | 93.0% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| algotune | 18 models · 18 scores | gpt-5.2-2025-12-11 | 2.05 | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| cad eval | 15 models · 15 scores | o3-2025-04-16 | 74.0% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| cl bench life | 14 models · 17 scores | gpt-5.5 | 22.2% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| the agent company | 14 models · 14 scores | DeepSeek-V3.2-Exp | 52.4% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| adversarial nli | 11 models · 11 scores | Phi-3-small-8k-instruct | 58.1% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| gdpval | 11 models · 11 scores | gpt-5.2-2025-12-11 | 49.7% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| posttrainbench | 11 models · 11 scores | claude-fable-5 | 41.8% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| rli | 11 models · 13 scores | claude-fable-5 | 16.1% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| exploitbench | 9 models · 10 scores | claude-mythos-preview-early | 73.8% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| frontierswe | 9 models · 9 scores | claude-fable-5-1✓ | 56.3% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| os world | 9 models · 20 scores | claude-sonnet-4-6 | 72.1 | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| osworld 2 | 9 models · 14 scores | claude-opus-5 | 31.4% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| spatialviz bench | 8 models · 8 scores | gemini-2.5-pro | 44.7% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| superglue | 8 models · 8 scores | T5-11B | 88.9% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| btf3 | 6 models · 8 scores | claude-sonnet-5 | 15.4% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| common sense qa 2 | 6 models · 6 scores | T5-11B | 67.8% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| mindcube | 5 models · 5 scores | gemma-3-12b-it | 46.7% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| proofbenchretired | 64 models · 64 scores | claude-fable-5-1✓ | 100.0% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
| fictionlivebenchretired | 41 models · 42 scores | o3-2025-04-16 | 100.0% | Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real. |
“Top scorer” is the highest number recorded in that column, not a leader in the Board's sense — the Board's leaders carry a date, a scaffold and a dispersion estimate, and these do not. A model can appear more than once in a series at different scaffolds; the figure shown is its best. Scores are in each series' own units — 53 of these are rates and render as percentages; the other 11 are Elo, points or time horizons and are shown as the source reports them, with no unit claimed. Series data from Epoch AI, used under CC-BY 4.0.