Which is best?
A benchmark is a fixed set of questions or tasks, put to every model the same way — graduate physics problems, real bugs from open-source projects, chess positions. The percentage beside a model is the share of that particular test it got right, and nothing else: 94% on one test and 48% on another are answers to two different questions, not a ranking.
There is no overall winner here, and that is the finding rather than a gap. Nobody has calibrated these tests against each other, so a model can top one and sit mid-table on the next — which is information about what it is good at, not a contradiction to be averaged away into a single rank.
Each test below is kept as its own list, and a test appears here only if whoever ran it published three things: the date it was run, an error bar, and a route to the run itself. Most published benchmark numbers have none of the three.
Sixty-four further tests are collected and deliberately not published here — what else we measure, and why it is not on the Board.
Two things travel with every score here. The date is the day the test was run, never the day the model came out — the gap between those two is where a model quietly acquires the answers. And the setting is named, because the same model on the same test scores differently depending on how long it was allowed to think and how many attempts it got: on every test here that swing is wider than the gap between the top two models, and on one of them it is eighty-eight times wider. A score without its setting is a measurement of the setting.
chess puzzlesEpoch AI
Games, played through to a result.
- 172.0% (max)±4.5gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
- 264.0% (promax)±4.8gpt-5.6-solOpenAI7.0%–64.0% across 4 settingstests disagreeevaluated 2026-08-07
- 364.0% (xhigh)±4.8gpt-5.5-pro-pre-releaseOpenAItests agreeevaluated 2026-05-01
- 461.0% (high)±4.9gemini-3.8-flashGoogle DeepMindtests agreeevaluated 2026-09-02
- 558.6% (xhigh)±5.0gpt-5.4-pro-2026-03-05OpenAItests agreeevaluated 2026-03-19
- 655.0%±5.0gemini-3.1-pro-previewGoogle DeepMind49.0%–55.0% across 2 settingstests disagreeevaluated 2026-08-06
- 754.0% (max)±5.0gpt-5.6-terraOpenAI5.0%–54.0% across 3 settingstests agreeevaluated 2026-08-07
- 854.0% (xhigh)±5.0gpt-5.5-pre-releaseOpenAItests agreeevaluated 2026-04-24
- 950.0% (high)±5.0gemini-3.5-flashGoogle DeepMind43.0%–50.0% across 3 settingstests agreeevaluated 2026-08-06
- 1049.0% (xhigh)±5.0gpt-5.2-2025-12-11OpenAI4.0%–49.0% across 5 settingstests agreeevaluated 2026-07-13
- 1147.0% (max)±5.0deepseek-v4-pro-0813DeepSeektests agreeevaluated 2026-08-18
- 1247.0% (high)±5.0gemini-3.7-flashGoogle DeepMindtests agreeevaluated 2026-08-14
ebr benchEpoch AI
- 176.2% (max)±8.4gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
- 257.1% (max)±6.0claude-fable-5-1Anthropictests agreeevaluated 2026-09-08
- 345.7% (max)±3.6claude-opus-5Anthropictests agreeevaluated 2026-09-05
- 444.8% (max)±4.2gpt-5.6-solOpenAItests disagreeevaluated 2026-09-04
- 539.5% (max)±2.0claude-fable-5Anthropictests agreeevaluated 2026-07-17
- 634.3% (xhigh)±2.0gpt-5.5OpenAItests disagreeevaluated 2026-07-27
- 730.5% (xhigh)±1.6grok-4.6xAItests agreeevaluated 2026-08-18
- 828.6% (max)±1.4claude-opus-4-8Anthropictests agreeevaluated 2026-08-07
- 925.4% (xhigh)gpt-5.4-2026-03-05OpenAItests agreeevaluated 2026-06-25
- 1023.0% (xhigh)gpt-5.2-2025-12-11OpenAItests agreeevaluated 2026-06-26
- 1119.0% (max)claude-opus-4-7Anthropictests agreeevaluated 2026-06-30
- 1214.3%gemini-3.1-pro-previewGoogle DeepMindtests disagreeevaluated 2026-06-25
frontiermathEpoch AI
Unpublished research-level maths problems, so no model can have been trained on them.
- 152.4% (high)±2.9gpt-5.5-pro-pre-releaseOpenAI51.0%–52.4% across 2 settingstests agreeevaluated 2026-04-23
- 251.7% (xhigh)±2.9gpt-5.5-pre-releaseOpenAItests agreeevaluated 2026-04-23
- 350.0% (xhigh)±2.9gpt-5.4-pro-2026-03-05OpenAItests agreeevaluated 2026-03-06
- 447.6% (xhigh)±2.9gpt-5.4-2026-03-05OpenAItests agreeevaluated 2026-03-06
- 547.2% (max)±2.9claude-opus-4-8Anthropictests agreeevaluated 2026-06-08
- 643.8% (xhigh)±2.9claude-opus-4-7Anthropictests agreeevaluated 2026-04-17
- 740.7% (max)±2.9claude-opus-4-6Anthropic38.3%–40.7% across 4 settingstests agreeevaluated 2026-02-12
- 840.7% (xhigh)±2.9gpt-5.2-2025-12-11OpenAI26.6%–40.7% across 4 settingstests agreeevaluated 2025-12-13
- 939.0%±2.9muse-sparkMeta AItests agreeevaluated 2026-04-08
- 1039.0%±2.9kimi-k2.6Moonshottests agreeevaluated 2026-05-07
- 1139.0% (high)±2.9gemini-3.5-flashGoogle DeepMindtests agreeevaluated 2026-05-22
- 1237.6%±2.8gemini-3-pro-previewGoogle DeepMindtests agreeevaluated 2025-11-21
frontiermath erdosEpoch AI
- 12.9% (max)±2.1gpt-6-astraOpenAItests disagreeevaluated 2026-08-28
- 20.0% (max)±0.0gpt-5.6-solOpenAItests disagreeevaluated 2026-08-28
- 30.0% (xhigh)±0.0gpt-5.5OpenAItests disagreeevaluated 2026-08-28
- 40.0% (max)±0.0claude-fable-5Anthropictests agreeevaluated 2026-08-28
- 50.0% (max)±0.0claude-fable-5-1Anthropictests agreeevaluated 2026-09-01
frontiermath tier 4Epoch AI
The hardest tier of FrontierMath's unpublished problems — set for research mathematicians.
- 147.9%±7.2gdm-ai-co-mathematicianGoogle DeepMindtoo few tests to checkevaluated 2026-05-08
- 239.6% (high)±7.1gpt-5.5-pro-pre-releaseOpenAI39.6%–39.6% across 2 settingstests agreeevaluated 2026-04-23
- 337.5%±7.0gpt-5.4-pro-2026-03-05-web-appOpenAItoo few tests to checkevaluated 2026-03-06
- 435.4% (xhigh)±6.9gpt-5.5-pre-releaseOpenAItests agreeevaluated 2026-04-23
- 531.3%±6.7gpt-5.2-pro-2025-12-11-webappOpenAItoo few tests to checkevaluated 2025-12-24
- 631.3% (max)±6.8claude-opus-4-8Anthropictests agreeevaluated 2026-06-08
- 727.1% (xhigh)±6.4gpt-5.4-2026-03-05OpenAItests agreeevaluated 2026-03-06
- 822.9% (xhigh)±6.1claude-opus-4-7Anthropictests agreeevaluated 2026-04-17
- 922.9% (max)±6.1claude-opus-4-6Anthropic14.6%–22.9% across 4 settingstests agreeevaluated 2026-02-12
- 1018.8% (high)±5.6gpt-5.2-2025-12-11OpenAI6.3%–18.8% across 4 settingstests agreeevaluated 2025-12-14
- 1118.8%±5.7gemini-3-pro-previewGoogle DeepMindtests agreeevaluated 2025-11-21
- 1216.7%±5.4gemini-3.1-pro-previewGoogle DeepMindtests disagreeevaluated 2026-02-19
frontiermath tier 4 v2Epoch AI
- 197.6% (high)±2.4gpt-6-astraOpenAI82.9%–97.6% across 6 settingstests disagreeevaluated 2026-08-30
- 290.2% (max)±4.6claude-fable-5Anthropictests agreeevaluated 2026-06-09
- 387.8% (max)±5.2claude-fable-5-1Anthropictests agreeevaluated 2026-09-01
- 482.9% (max)±5.9gpt-5.6-solOpenAI80.5%–82.9% across 2 settingstests disagreeevaluated 2026-07-09
- 578.0% (xhigh)±6.5gpt-5.5-proOpenAItoo few tests to checkevaluated 2026-06-12
- 675.6%±6.7gdm-ai-co-mathematicianGoogle DeepMindtoo few tests to checkevaluated 2026-06-12
- 773.2% (max)±7.0claude-opus-5Anthropictests agreeevaluated 2026-07-24
- 872.5% (xhigh)±7.1gpt-5.5OpenAItests disagreeevaluated 2026-06-11
- 970.7% (max)±7.2gpt-5.6-terraOpenAItests agreeevaluated 2026-07-09
- 1061.0% (max)±7.7gpt-5.6-lunaOpenAItests agreeevaluated 2026-07-09
- 1158.5% (xhigh)±7.8gpt-5.4-pro-2026-03-05OpenAItests agreeevaluated 2026-06-13
- 1256.1% (max)±7.8claude-opus-4-8Anthropictests agreeevaluated 2026-06-10
frontiermath tiers 1 3 v2Epoch AI
- 193.7% (max)±1.4gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
- 290.2% (max)±1.8claude-fable-5-1Anthropictests agreeevaluated 2026-09-01
- 389.1% (max)±1.8gpt-5.6-solOpenAItests disagreeevaluated 2026-07-09
- 487.7% (xhigh)±1.9gpt-5.5-proOpenAItoo few tests to checkevaluated 2026-06-12
- 587.0% (max)±2.0claude-fable-5Anthropictests agreeevaluated 2026-06-09
- 686.0% (max)±2.1gpt-5.6-terraOpenAItests agreeevaluated 2026-07-09
- 785.6% (max)±2.1claude-opus-5Anthropictests agreeevaluated 2026-07-24
- 885.3% (xhigh)±2.1gpt-5.5OpenAItests disagreeevaluated 2026-06-11
- 982.5% (xhigh)±2.3gpt-5.4-pro-2026-03-05OpenAItests agreeevaluated 2026-06-13
- 1082.1% (max)±2.3gpt-5.6-lunaOpenAI39.6%–82.1% across 3 settingstests agreeevaluated 2026-08-29
- 1180.0% (max)±2.4claude-opus-4-8Anthropictests agreeevaluated 2026-06-10
- 1278.6% (xhigh)±2.4gpt-5.4-2026-03-05OpenAItests agreeevaluated 2026-06-11
gpqa diamondEpoch AI
Graduate-level science questions, written so the answer cannot simply be looked up.
- 195.8% (max)±1.4gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
- 295.4% (high)±1.4gemini-3.8-flashGoogle DeepMindtests agreeevaluated 2026-09-02
- 394.8% (high)±1.3gemini-3.7-flashGoogle DeepMindtests agreeevaluated 2026-08-14
- 494.6% (xhigh)±1.6gpt-5.4-pro-2026-03-05OpenAItests agreeevaluated 2026-03-20
- 594.4% (high)±1.6gemini-3.1-pro-previewGoogle DeepMind94.1%–94.4% across 2 settingstests disagreeevaluated 2026-08-06
- 694.1% (high)±1.4gemini-3.6-flashGoogle DeepMind85.9%–94.1% across 3 settingstests agreeevaluated 2026-08-07
- 794.0% (high)±1.4grok-4.6xAI93.2%–94.0% across 2 settingstests agreeevaluated 2026-08-14
- 894.0% (xhigh)±1.5gpt-5.5-pre-releaseOpenAItests agreeevaluated 2026-04-24
- 993.9% (xhigh)±1.6gpt-5.5-pro-pre-releaseOpenAItests agreeevaluated 2026-04-24
- 1093.9% (max)±1.5claude-opus-5Anthropic87.9%–93.9% across 3 settingstests agreeevaluated 2026-08-06
- 1193.5% (max)±1.6gpt-5.6-solOpenAI82.8%–93.5% across 3 settingstests disagreeevaluated 2026-08-07
- 1293.4% (high)±1.4grok-4.5xAItests agreeevaluated 2026-07-08
mirrorcodeEpoch AI
Real tasks to carry out, rather than questions to answer.
- 173.3% (high)claude-fable-5-1Anthropictests agreeevaluated 2026-09-10
- 263.9% (high)±10.4claude-fable-5Anthropictests agreeevaluated 2026-08-10
- 346.7% (high)gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
- 431.1% (high)±9.0claude-opus-4-7Anthropictests agreeevaluated 2026-08-12
- 520.0% (high)±9.2gpt-5.6-solOpenAItests disagreeevaluated 2026-08-10
- 615.6% (high)±7.7gpt-5.4-2026-03-05OpenAItests agreeevaluated 2026-08-10
- 710.0% (high)±6.0gpt-5.5OpenAItests disagreeevaluated 2026-08-10
- 88.9% (high)±4.6gemini-3.1-pro-previewGoogle DeepMindtests disagreeevaluated 2026-08-10
mystery game puzzlesEpoch AI
Games, played through to a result.
- 184.0% (max)±3.7gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
- 259.0% (max)±4.9claude-opus-5Anthropic37.0%–59.0% across 2 settingstests agreeevaluated 2026-08-06
- 358.0% (max)±5.0gpt-5.6-solOpenAI26.0%–58.0% across 3 settingstests disagreeevaluated 2026-08-27
- 458.0% (max)±5.0claude-fable-5-1Anthropictests agreeevaluated 2026-09-01
- 556.0% (xhigh)±5.0gpt-5.5OpenAI18.0%–56.0% across 4 settingstests disagreeevaluated 2026-08-28
- 652.0% (max)±5.0claude-fable-5Anthropictests agreeevaluated 2026-07-31
- 747.0% (high)±5.0gemini-3.8-flashGoogle DeepMindtests agreeevaluated 2026-09-02
- 843.0% (max)±5.0deepseek-v4-pro-0813DeepSeektests agreeevaluated 2026-08-19
- 938.0% (xhigh)±4.9qwen3.8-maxAlibabatests agreeevaluated 2026-08-05
- 1037.0% (high)±4.9gemini-3.7-flashGoogle DeepMindtests agreeevaluated 2026-08-14
- 1137.0% (xhigh)±4.9gpt-5.4-2026-03-05OpenAI16.0%–37.0% across 4 settingstests agreeevaluated 2026-08-30
- 1236.0% (max)±4.8claude-opus-4-8Anthropic31.0%–36.0% across 2 settingstests agreeevaluated 2026-07-26
simpleqa verifiedEpoch AI
Short factual questions with one checkable answer — a test of whether a model knows or invents.
- 175.6% (max)±1.4gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
- 273.5% (high)±1.4gemini-3.1-pro-previewGoogle DeepMindtests disagreeevaluated 2026-08-10
- 370.8% (max)±1.4claude-fable-5-1Anthropictests agreeevaluated 2026-09-01
- 470.7% (xhigh)±1.4claude-fable-5Anthropictests agreeevaluated 2026-08-10
- 569.7% (max)±1.5gpt-5.6-solOpenAItests disagreeevaluated 2026-08-10
- 669.7% (high)±1.5gemini-3.8-flashGoogle DeepMindtests agreeevaluated 2026-09-02
- 769.2% (high)±1.5gemini-3.7-flashGoogle DeepMindtests agreeevaluated 2026-08-27
- 866.8% (high)±1.5gemini-3-flash-previewGoogle DeepMindtests agreeevaluated 2026-08-27
- 966.2% (high)±1.5gemini-3.6-flashGoogle DeepMindtests agreeevaluated 2026-08-27
- 1066.2% (high)±1.5gemini-3.5-flashGoogle DeepMindtests agreeevaluated 2026-08-27
- 1163.0% (xhigh)±1.5gpt-5.5OpenAItests disagreeevaluated 2026-08-27
- 1260.3% (xhigh)±1.6muse-spark-1.2Meta AItoo few tests to checkevaluated 2026-08-27
swe bench verifiedEpoch AI
Real bugs from open-source projects: the fix has to make the project's own tests pass.
- 183.5% (max)±1.7claude-opus-4-7Anthropictests agreeevaluated 2026-04-20
- 280.6% (xhigh)±1.8gpt-5.5-pre-releaseOpenAItests agreeevaluated 2026-04-24
- 379.3% (high)±1.8gemini-3.5-flashGoogle DeepMindtests agreeevaluated 2026-06-01
- 478.7%±1.9claude-opus-4-6Anthropic75.6%–78.7% across 2 settingstests agreeevaluated 2026-02-18
- 578.7% (max)±1.9glm-5.2Z.ai (Zhipu AI)tests agreeevaluated 2026-06-25
- 677.6% (max)±1.9deepseek-v4-proDeepSeektests agreeevaluated 2026-06-18
- 777.3%±1.9qwen3.7-maxAlibabatests agreeevaluated 2026-06-18
- 876.9% (high)±1.9gpt-5.4-2026-03-05OpenAItests agreeevaluated 2026-03-06
- 976.7%±1.9qwen3.6-max-previewAlibabatests agreeevaluated 2026-05-28
- 1076.7%±1.9claude-opus-4-5-20251101Anthropictests agreeevaluated 2026-02-05
- 1176.7%±1.9kimi-k2.6Moonshottests agreeevaluated 2026-05-08
- 1275.6%±2.0gemini-3.1-pro-preview-customtoolsGoogle DeepMindtoo few tests to checkevaluated 2026-02-24
math level 5Epoch AIsaturating
The hardest tier of a competition maths set that has been public for years.
Saturating: 20% of all scores are above 90%. Treat small differences near the top as noise — the benchmark is running out of room before the models run out of ability.
- 198.1% (high)±0.3gpt-5-2025-08-07OpenAI97.9%–98.1% across 2 settingstests agreeevaluated 2025-10-29
- 297.8% (high)±0.3gpt-5-mini-2025-08-07OpenAI96.8%–97.8% across 2 settingstests agreeevaluated 2025-10-30
- 397.8% (high)±0.3o4-mini-2025-04-16OpenAItests disagreeevaluated 2025-04-16
- 497.8% (high)±0.3o3-2025-04-16OpenAItests agreeevaluated 2025-04-16
- 597.7% (32K)±0.4claude-sonnet-4-5-20250929Anthropictests agreeevaluated 2025-10-21
- 697.1%±0.4qwen3-max-2025-09-23Alibabatests disagreeevaluated 2025-10-09
- 796.6%±0.4DeepSeek-R1-0528DeepSeektests agreeevaluated 2025-05-29
- 896.5% (high)±0.4o3-mini-2025-01-31OpenAI95.2%–96.5% across 2 settingstests disagreeevaluated 2025-02-13
- 996.4% (32K)±0.5claude-haiku-4-5-20251001Anthropic86.9%–96.4% across 2 settingstests disagreeevaluated 2025-10-22
- 1095.9%±0.4gemini-2.5-pro-preview-05-06Google DeepMindtoo few tests to checkevaluated 2025-05-08
- 1195.6%±0.4gemini-2.5-pro-preview-03-25Google DeepMindtoo few tests to checkevaluated 2025-05-07
- 1295.2% (medium)±0.5gpt-5-nano-2025-08-07OpenAI94.9%–95.2% across 2 settingstests disagreeevaluated 2025-08-20
otis mock aime 2024 2025Epoch AIretired
Mock papers for AIME, the exam that qualifies US high-school students for the maths olympiad.
Retired as an instrument: some model has scored 100.0% — a perfect result — so this no longer separates the models at the top. The series is kept because when a benchmark stopped working is itself a finding, but nothing new is added to it.
- 1100.0% (max)±0.0gpt-5.6-solOpenAI68.9%–100.0% across 3 settingstests disagreeevaluated 2026-08-07
- 2100.0% (high)±0.0claude-fable-5Anthropic97.8%–100.0% across 3 settingstests agreeevaluated 2026-08-06
- 3100.0% (xhigh)±0.0gpt-5.5-pro-pre-releaseOpenAItests agreeevaluated 2026-04-24
- 4100.0% (xhigh)±0.0gpt-5.5-pre-releaseOpenAItests agreeevaluated 2026-04-24
- 5100.0% (max)±0.0claude-fable-5-1Anthropictests agreeevaluated 2026-09-01
- 6100.0% (max)±0.0gpt-6-astraOpenAItests disagreeevaluated 2026-08-30
- 7100.0% (xhigh)±0.0qwen3.8-max-0902Alibabatests agreeevaluated 2026-09-02
- 899.7% (max)±0.3gpt-5.6-terraOpenAI53.3%–99.7% across 3 settingstests agreeevaluated 2026-08-07
- 999.4% (xhigh)±0.4qwen3.8-maxAlibabatests agreeevaluated 2026-08-04
- 1099.2% (xhigh)±0.5grok-4.6xAI97.8%–99.2% across 2 settingstests agreeevaluated 2026-08-14
- 1198.9% (max)±1.1claude-opus-5Anthropic93.3%–98.9% across 3 settingstests agreeevaluated 2026-08-06
- 1298.9% (high)±0.9gemini-3.8-flashGoogle DeepMindtests agreeevaluated 2026-09-02
Benchmark data: Epoch AI, 'AI Benchmarking Hub'. https://epoch.ai/benchmarks. CC-BY-4.0.Datum computes the ranking, the scaffold range and the agreement test; the measurements are Epoch's.