What else we measure

Datum ingests every benchmark series in Epoch's archive. The Board publishes 14. These are the 66 it does not, and why.

A score with no evaluation date could have been run at any point, including after the questions leaked into a model's training data — and the gap between release and evaluation is exactly where contamination lives. A score with no error bar cannot tell you whether the distance between two models is real or noise. Series missing either are held, not published.

They still earn their place. If a model leads a Board house and tops a stack of series measured independently, that agreement is evidence. If it leads on the Board and nowhere else, that is worth knowing too. So this page counts agreement — and never turns it into a rank, because none of these series is measured well enough to support one.

17 of 66

of these series are topped by a model that also leads a house on the Board. That is a count of agreement, not a score — it has no threshold and no pass mark, because a threshold would be a verdict these numbers cannot carry.

Every series we hold

SeriesCoverageTop scorerBest scoreWhy it is not on the Board
dtbenchsaturating160 models · 161 scoresclaude-fable-598.4%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
mmlu136 models · 136 scoresgpt-4o-2024-11-2088.1%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
weirdml135 models · 169 scoresclaude-fable-5-192.9%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
critpt123 models · 177 scoresgpt-5.6-sol32.3%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
lmca122 models · 123 scoresclaude-opus-563.3Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
scicode113 models · 165 scoresclaude-fable-5-162.0%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
ale bench102 models · 112 scoresgpt-5.6-sol2177Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
webdev arenaarena94 models · 120 scoresgpt-6-astra1797Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
gsm8k93 models · 93 scoresDeepSeek-Coder-V2-Instruct94.5%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
simplebench87 models · 102 scoresclaude-fable-581.9%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
arc agisaturating83 models · 213 scoresgpt-6-astra98.5%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
arc agi 280 models · 203 scoresgpt-6-astra95.0%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
wino grande80 models · 80 scoresLlama-3.1-405B89.2%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
forecastbench78 models · 81 scoreso3-2025-04-1662.5Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
arc ai277 models · 77 scoresLlama-3.1-405B95.3%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
bool q77 models · 77 scoresT5-11B91.2%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
hella swag76 models · 76 scoresgpt-4-32k-031495.3%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
piqa60 models · 60 scoresgpt-4o-mini-2024-07-1888.7%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
aider polyglot58 models · 71 scoresgpt-5-2025-08-0788Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
vending bench 255 models · 60 scoresclaude-opus-511182Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
apex agents54 models · 65 scoresgpt-6-astra62.4%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
bbh50 models · 50 scoresgemini-1.5-pro-00189.2%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
live bench50 models · 52 scoresgemini-2.5-pro-exp-03-2582.35Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
video mme50 models · 50 scoresvideo-SALMONN-2plus79.7%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
terminalbench45 models · 146 scoresgpt-5.584.7%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
lech mazur writing44 models · 49 scoresgpt-5-2025-08-078.6Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
metr time horizons43 models · 50 scoresclaude-mythos-preview-early1045Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
hle42 models · 46 scoresclaude-fable-5-146.5%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
open book qa42 models · 42 scoresPhi-3-mini-4k-instruct88.0%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
enigma eval41 models · 46 scoresclaude-fable-539.3%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
trivia qa39 models · 39 scoresLlama-2-70b-hf87.6%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
balrog32 models · 35 scoresgemini-3-pro-preview58.1%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
frontiercode30 models · 30 scoresclaude-fable-553.5%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
vpct28 models · 38 scoresgemini-3-pro-preview91.0%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
geobench27 models · 32 scoresgemini-3-flash-preview4333Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
gso27 models · 27 scoresclaude-opus-4-847.1%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
deepswe26 models · 26 scoresgpt-6-astra74.1%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
lambada26 models · 26 scoresMegatron-Turing NLG 530B87.2%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
deepresearchbench25 models · 41 scoresclaude-opus-4-655.3%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
gdp pdf25 models · 39 scoresgpt-5.6-sol30.7%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
blueprint bench 224 models · 24 scoresclaude-fable-5-141.9%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
gbaeval23 models · 23 scoresclaude-opus-579.6%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
science qasaturating23 models · 23 scoresPhi-3.5-vision-instruct91.3%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
surface evolver bench23 models · 26 scoreskimi-k395.0%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
cursorbench22 models · 74 scoresclaude-fable-5-173.4%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
cl bench21 models · 23 scoresgpt-5.4-2026-03-0527.9%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
cybench21 models · 22 scoresclaude-opus-4-693.0%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
algotune18 models · 18 scoresgpt-5.2-2025-12-112.05Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
cad eval15 models · 15 scoreso3-2025-04-1674.0%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
cl bench life14 models · 17 scoresgpt-5.522.2%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
the agent company14 models · 14 scoresDeepSeek-V3.2-Exp52.4%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
adversarial nli11 models · 11 scoresPhi-3-small-8k-instruct58.1%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
gdpval11 models · 11 scoresgpt-5.2-2025-12-1149.7%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
posttrainbench11 models · 11 scoresclaude-fable-541.8%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
rli11 models · 13 scoresclaude-fable-516.1%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
exploitbench9 models · 10 scoresclaude-mythos-preview-early73.8%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
frontierswe9 models · 9 scoresclaude-fable-5-156.3%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
os world9 models · 20 scoresclaude-sonnet-4-672.1Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
osworld 29 models · 14 scoresclaude-opus-531.4%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
spatialviz bench8 models · 8 scoresgemini-2.5-pro44.7%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
superglue8 models · 8 scoresT5-11B88.9%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
btf36 models · 8 scoresclaude-sonnet-515.4%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
common sense qa 26 models · 6 scoresT5-11B67.8%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
mindcube5 models · 5 scoresgemma-3-12b-it46.7%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
proofbenchretired64 models · 64 scoresclaude-fable-5-1100.0%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.
fictionlivebenchretired41 models · 42 scoreso3-2025-04-16100.0%Published with no evaluation date and no error bar, so we cannot say when they were run or whether a gap between two models is real.

“Top scorer” is the highest number recorded in that column, not a leader in the Board's sense — the Board's leaders carry a date, a scaffold and a dispersion estimate, and these do not. A model can appear more than once in a series at different scaffolds; the figure shown is its best. Scores are in each series' own units — 53 of these are rates and render as percentages; the other 11 are Elo, points or time horizons and are shown as the source reports them, with no unit claimed. Series data from Epoch AI, used under CC-BY 4.0.