Which is best?

A benchmark is a fixed set of questions or tasks, put to every model the same way — graduate physics problems, real bugs from open-source projects, chess positions. The percentage beside a model is the share of that particular test it got right, and nothing else: 94% on one test and 48% on another are answers to two different questions, not a ranking.

There is no overall winner here, and that is the finding rather than a gap. Nobody has calibrated these tests against each other, so a model can top one and sit mid-table on the next — which is information about what it is good at, not a contradiction to be averaged away into a single rank.

Each test below is kept as its own list, and a test appears here only if whoever ran it published three things: the date it was run, an error bar, and a route to the run itself. Most published benchmark numbers have none of the three.

Sixty-four further tests are collected and deliberately not published here — what else we measure, and why it is not on the Board.

Two things travel with every score here. The date is the day the test was run, never the day the model came out — the gap between those two is where a model quietly acquires the answers. And the setting is named, because the same model on the same test scores differently depending on how long it was allowed to think and how many attempts it got: on every test here that swing is wider than the gap between the top two models, and on one of them it is eighty-eight times wider. A score without its setting is a measurement of the setting.

chess puzzlesEpoch AI

Games, played through to a result.

  • 172.0% (max)±4.5gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
  • 264.0% (promax)±4.8gpt-5.6-solOpenAI7.0%–64.0% across 4 settingstests disagreeevaluated 2026-08-07
  • 364.0% (xhigh)±4.8gpt-5.5-pro-pre-releaseOpenAItests agreeevaluated 2026-05-01
  • 461.0% (high)±4.9gemini-3.8-flashGoogle DeepMindtests agreeevaluated 2026-09-02
  • 558.6% (xhigh)±5.0gpt-5.4-pro-2026-03-05OpenAItests agreeevaluated 2026-03-19
  • 655.0%±5.0gemini-3.1-pro-previewGoogle DeepMind49.0%–55.0% across 2 settingstests disagreeevaluated 2026-08-06
  • 754.0% (max)±5.0gpt-5.6-terraOpenAI5.0%–54.0% across 3 settingstests agreeevaluated 2026-08-07
  • 854.0% (xhigh)±5.0gpt-5.5-pre-releaseOpenAItests agreeevaluated 2026-04-24
  • 950.0% (high)±5.0gemini-3.5-flashGoogle DeepMind43.0%–50.0% across 3 settingstests agreeevaluated 2026-08-06
  • 1049.0% (xhigh)±5.0gpt-5.2-2025-12-11OpenAI4.0%–49.0% across 5 settingstests agreeevaluated 2026-07-13
  • 1147.0% (max)±5.0deepseek-v4-pro-0813DeepSeektests agreeevaluated 2026-08-18
  • 1247.0% (high)±5.0gemini-3.7-flashGoogle DeepMindtests agreeevaluated 2026-08-14

ebr benchEpoch AI

frontiermathEpoch AI

Unpublished research-level maths problems, so no model can have been trained on them.

frontiermath erdosEpoch AI

  • 12.9% (max)±2.1gpt-6-astraOpenAItests disagreeevaluated 2026-08-28
  • 20.0% (max)±0.0gpt-5.6-solOpenAItests disagreeevaluated 2026-08-28
  • 30.0% (xhigh)±0.0gpt-5.5OpenAItests disagreeevaluated 2026-08-28
  • 40.0% (max)±0.0claude-fable-5Anthropictests agreeevaluated 2026-08-28
  • 50.0% (max)±0.0claude-fable-5-1Anthropictests agreeevaluated 2026-09-01

frontiermath tier 4Epoch AI

The hardest tier of FrontierMath's unpublished problems — set for research mathematicians.

frontiermath tier 4 v2Epoch AI

  • 197.6% (high)±2.4gpt-6-astraOpenAI82.9%–97.6% across 6 settingstests disagreeevaluated 2026-08-30
  • 290.2% (max)±4.6claude-fable-5Anthropictests agreeevaluated 2026-06-09
  • 387.8% (max)±5.2claude-fable-5-1Anthropictests agreeevaluated 2026-09-01
  • 482.9% (max)±5.9gpt-5.6-solOpenAI80.5%–82.9% across 2 settingstests disagreeevaluated 2026-07-09
  • 578.0% (xhigh)±6.5gpt-5.5-proOpenAItoo few tests to checkevaluated 2026-06-12
  • 675.6%±6.7gdm-ai-co-mathematicianGoogle DeepMindtoo few tests to checkevaluated 2026-06-12
  • 773.2% (max)±7.0claude-opus-5Anthropictests agreeevaluated 2026-07-24
  • 872.5% (xhigh)±7.1gpt-5.5OpenAItests disagreeevaluated 2026-06-11
  • 970.7% (max)±7.2gpt-5.6-terraOpenAItests agreeevaluated 2026-07-09
  • 1061.0% (max)±7.7gpt-5.6-lunaOpenAItests agreeevaluated 2026-07-09
  • 1158.5% (xhigh)±7.8gpt-5.4-pro-2026-03-05OpenAItests agreeevaluated 2026-06-13
  • 1256.1% (max)±7.8claude-opus-4-8Anthropictests agreeevaluated 2026-06-10

frontiermath tiers 1 3 v2Epoch AI

  • 193.7% (max)±1.4gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
  • 290.2% (max)±1.8claude-fable-5-1Anthropictests agreeevaluated 2026-09-01
  • 389.1% (max)±1.8gpt-5.6-solOpenAItests disagreeevaluated 2026-07-09
  • 487.7% (xhigh)±1.9gpt-5.5-proOpenAItoo few tests to checkevaluated 2026-06-12
  • 587.0% (max)±2.0claude-fable-5Anthropictests agreeevaluated 2026-06-09
  • 686.0% (max)±2.1gpt-5.6-terraOpenAItests agreeevaluated 2026-07-09
  • 785.6% (max)±2.1claude-opus-5Anthropictests agreeevaluated 2026-07-24
  • 885.3% (xhigh)±2.1gpt-5.5OpenAItests disagreeevaluated 2026-06-11
  • 982.5% (xhigh)±2.3gpt-5.4-pro-2026-03-05OpenAItests agreeevaluated 2026-06-13
  • 1082.1% (max)±2.3gpt-5.6-lunaOpenAI39.6%–82.1% across 3 settingstests agreeevaluated 2026-08-29
  • 1180.0% (max)±2.4claude-opus-4-8Anthropictests agreeevaluated 2026-06-10
  • 1278.6% (xhigh)±2.4gpt-5.4-2026-03-05OpenAItests agreeevaluated 2026-06-11

gpqa diamondEpoch AI

Graduate-level science questions, written so the answer cannot simply be looked up.

  • 195.8% (max)±1.4gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
  • 295.4% (high)±1.4gemini-3.8-flashGoogle DeepMindtests agreeevaluated 2026-09-02
  • 394.8% (high)±1.3gemini-3.7-flashGoogle DeepMindtests agreeevaluated 2026-08-14
  • 494.6% (xhigh)±1.6gpt-5.4-pro-2026-03-05OpenAItests agreeevaluated 2026-03-20
  • 594.4% (high)±1.6gemini-3.1-pro-previewGoogle DeepMind94.1%–94.4% across 2 settingstests disagreeevaluated 2026-08-06
  • 694.1% (high)±1.4gemini-3.6-flashGoogle DeepMind85.9%–94.1% across 3 settingstests agreeevaluated 2026-08-07
  • 794.0% (high)±1.4grok-4.6xAI93.2%–94.0% across 2 settingstests agreeevaluated 2026-08-14
  • 894.0% (xhigh)±1.5gpt-5.5-pre-releaseOpenAItests agreeevaluated 2026-04-24
  • 993.9% (xhigh)±1.6gpt-5.5-pro-pre-releaseOpenAItests agreeevaluated 2026-04-24
  • 1093.9% (max)±1.5claude-opus-5Anthropic87.9%–93.9% across 3 settingstests agreeevaluated 2026-08-06
  • 1193.5% (max)±1.6gpt-5.6-solOpenAI82.8%–93.5% across 3 settingstests disagreeevaluated 2026-08-07
  • 1293.4% (high)±1.4grok-4.5xAItests agreeevaluated 2026-07-08

mirrorcodeEpoch AI

Real tasks to carry out, rather than questions to answer.

  • 173.3% (high)claude-fable-5-1Anthropictests agreeevaluated 2026-09-10
  • 263.9% (high)±10.4claude-fable-5Anthropictests agreeevaluated 2026-08-10
  • 346.7% (high)gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
  • 431.1% (high)±9.0claude-opus-4-7Anthropictests agreeevaluated 2026-08-12
  • 520.0% (high)±9.2gpt-5.6-solOpenAItests disagreeevaluated 2026-08-10
  • 615.6% (high)±7.7gpt-5.4-2026-03-05OpenAItests agreeevaluated 2026-08-10
  • 710.0% (high)±6.0gpt-5.5OpenAItests disagreeevaluated 2026-08-10
  • 88.9% (high)±4.6gemini-3.1-pro-previewGoogle DeepMindtests disagreeevaluated 2026-08-10

mystery game puzzlesEpoch AI

Games, played through to a result.

  • 184.0% (max)±3.7gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
  • 259.0% (max)±4.9claude-opus-5Anthropic37.0%–59.0% across 2 settingstests agreeevaluated 2026-08-06
  • 358.0% (max)±5.0gpt-5.6-solOpenAI26.0%–58.0% across 3 settingstests disagreeevaluated 2026-08-27
  • 458.0% (max)±5.0claude-fable-5-1Anthropictests agreeevaluated 2026-09-01
  • 556.0% (xhigh)±5.0gpt-5.5OpenAI18.0%–56.0% across 4 settingstests disagreeevaluated 2026-08-28
  • 652.0% (max)±5.0claude-fable-5Anthropictests agreeevaluated 2026-07-31
  • 747.0% (high)±5.0gemini-3.8-flashGoogle DeepMindtests agreeevaluated 2026-09-02
  • 843.0% (max)±5.0deepseek-v4-pro-0813DeepSeektests agreeevaluated 2026-08-19
  • 938.0% (xhigh)±4.9qwen3.8-maxAlibabatests agreeevaluated 2026-08-05
  • 1037.0% (high)±4.9gemini-3.7-flashGoogle DeepMindtests agreeevaluated 2026-08-14
  • 1137.0% (xhigh)±4.9gpt-5.4-2026-03-05OpenAI16.0%–37.0% across 4 settingstests agreeevaluated 2026-08-30
  • 1236.0% (max)±4.8claude-opus-4-8Anthropic31.0%–36.0% across 2 settingstests agreeevaluated 2026-07-26

simpleqa verifiedEpoch AI

Short factual questions with one checkable answer — a test of whether a model knows or invents.

  • 175.6% (max)±1.4gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
  • 273.5% (high)±1.4gemini-3.1-pro-previewGoogle DeepMindtests disagreeevaluated 2026-08-10
  • 370.8% (max)±1.4claude-fable-5-1Anthropictests agreeevaluated 2026-09-01
  • 470.7% (xhigh)±1.4claude-fable-5Anthropictests agreeevaluated 2026-08-10
  • 569.7% (max)±1.5gpt-5.6-solOpenAItests disagreeevaluated 2026-08-10
  • 669.7% (high)±1.5gemini-3.8-flashGoogle DeepMindtests agreeevaluated 2026-09-02
  • 769.2% (high)±1.5gemini-3.7-flashGoogle DeepMindtests agreeevaluated 2026-08-27
  • 866.8% (high)±1.5gemini-3-flash-previewGoogle DeepMindtests agreeevaluated 2026-08-27
  • 966.2% (high)±1.5gemini-3.6-flashGoogle DeepMindtests agreeevaluated 2026-08-27
  • 1066.2% (high)±1.5gemini-3.5-flashGoogle DeepMindtests agreeevaluated 2026-08-27
  • 1163.0% (xhigh)±1.5gpt-5.5OpenAItests disagreeevaluated 2026-08-27
  • 1260.3% (xhigh)±1.6muse-spark-1.2Meta AItoo few tests to checkevaluated 2026-08-27

swe bench verifiedEpoch AI

Real bugs from open-source projects: the fix has to make the project's own tests pass.

math level 5Epoch AIsaturating

The hardest tier of a competition maths set that has been public for years.

Saturating: 20% of all scores are above 90%. Treat small differences near the top as noise — the benchmark is running out of room before the models run out of ability.

otis mock aime 2024 2025Epoch AIretired

Mock papers for AIME, the exam that qualifies US high-school students for the maths olympiad.

Retired as an instrument: some model has scored 100.0% — a perfect result — so this no longer separates the models at the top. The series is kept because when a benchmark stopped working is itself a finding, but nothing new is added to it.

  • 1100.0% (max)±0.0gpt-5.6-solOpenAI68.9%–100.0% across 3 settingstests disagreeevaluated 2026-08-07
  • 2100.0% (high)±0.0claude-fable-5Anthropic97.8%–100.0% across 3 settingstests agreeevaluated 2026-08-06
  • 3100.0% (xhigh)±0.0gpt-5.5-pro-pre-releaseOpenAItests agreeevaluated 2026-04-24
  • 4100.0% (xhigh)±0.0gpt-5.5-pre-releaseOpenAItests agreeevaluated 2026-04-24
  • 5100.0% (max)±0.0claude-fable-5-1Anthropictests agreeevaluated 2026-09-01
  • 6100.0% (max)±0.0gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
  • 7100.0% (xhigh)±0.0qwen3.8-max-0902Alibabatests agreeevaluated 2026-09-02
  • 899.7% (max)±0.3gpt-5.6-terraOpenAI53.3%–99.7% across 3 settingstests agreeevaluated 2026-08-07
  • 999.4% (xhigh)±0.4qwen3.8-maxAlibabatests agreeevaluated 2026-08-04
  • 1099.2% (xhigh)±0.5grok-4.6xAI97.8%–99.2% across 2 settingstests agreeevaluated 2026-08-14
  • 1198.9% (max)±1.1claude-opus-5Anthropic93.3%–98.9% across 3 settingstests agreeevaluated 2026-08-06
  • 1298.9% (high)±0.9gemini-3.8-flashGoogle DeepMindtests agreeevaluated 2026-09-02

Benchmark data: Epoch AI, 'AI Benchmarking Hub'. https://epoch.ai/benchmarks. CC-BY-4.0.Datum computes the ranking, the scaffold range and the agreement test; the measurements are Epoch's.