OpenAI

o3-2025-04-16

Made by OpenAI. Released 2025-04-16. Accessibility not recorded. We hold no price for this model: no seller we can read lists it any more, so what follows is what was measured, not what it cost.

What it scores, by house

BenchmarkHouseScoreEvaluatedScaffold
chess puzzlesEpoch AI0.3±0.052026-08-07high
chess puzzlesEpoch AI0.4±0.052026-08-07medium
chess puzzlesEpoch AI0.3±0.042026-07-15low
frontiermathEpoch AI0.1±0.022025-11-17low
frontiermathEpoch AI0.2±0.022025-11-16medium
frontiermathEpoch AI0.2±0.022025-11-16high
frontiermath tier 4Epoch AI0.0±0.022025-07-01high
frontiermath tiers 1 3 v2Epoch AI0.3±0.032026-08-27high
frontiermath tiers 1 3 v2Epoch AI0.2±0.022026-08-27low
frontiermath tiers 1 3 v2Epoch AI0.3±0.032026-08-27medium
gpqa diamondEpoch AI0.8±0.032026-08-07medium
gpqa diamondEpoch AI0.8±0.032026-07-15low
gpqa diamondEpoch AI0.8±0.022025-04-16high
math level 5near its ceilingEpoch AI1.0±0.002025-04-16high
mystery game puzzlesEpoch AI0.2±0.042026-08-27medium
mystery game puzzlesEpoch AI0.3±0.052026-08-27high
otis mock aime 2024 2025near its ceilingEpoch AI0.8±0.052026-08-07medium
otis mock aime 2024 2025near its ceilingEpoch AI0.6±0.072026-07-15low
otis mock aime 2024 2025near its ceilingEpoch AI0.8±0.042025-04-16high
simpleqa verifiedEpoch AI0.5±0.022026-08-27high
swe bench verifiedEpoch AI0.6±0.022026-02-12medium
aider polyglotcollected by Epoch AI76.9unknown
aider polyglotcollected by Epoch AI76.9medium
aider polyglotcollected by Epoch AI81.3high
ale benchcollected by Epoch AI934high
apex agentscollected by Epoch AI0.3high
arc aginear its ceilingcollected by Epoch AI0.5medium
arc aginear its ceilingcollected by Epoch AI0.4low
arc aginear its ceilingcollected by Epoch AI0.6high
arc agi 2collected by Epoch AI0.1high
arc agi 2collected by Epoch AI0.0low
arc agi 2collected by Epoch AI0.0medium
cad evalcollected by Epoch AI0.7medium
cl benchcollected by Epoch AI0.2high
critptcollected by Epoch AI0.0high
deepresearchbenchcollected by Epoch AI0.5medium
dtbenchnear its ceilingcollected by Epoch AI0.8high
enigma evalcollected by Epoch AI0.1high
enigma evalcollected by Epoch AI0.1medium
fictionlivebenchcollected by Epoch AI1.0medium
forecastbenchcollected by Epoch AI62.5unknown
gdpvalcollected by Epoch AI0.3medium
geobenchcollected by Epoch AI3930high
geobenchcollected by Epoch AI3858medium
gsocollected by Epoch AI0.1OpenHands
hlecollected by Epoch AI0.2medium
hlecollected by Epoch AI0.2high
lech mazur writingcollected by Epoch AI8.4medium
lmcacollected by Epoch AI39.7high
metr time horizonscollected by Epoch AI120unknown
metr time horizonscollected by Epoch AI91.3medium
os worldcollected by Epoch AI9.1o3 (15 steps)
os worldcollected by Epoch AI23.0o3 (100 steps)
os worldcollected by Epoch AI17.2o3 (50 steps)
simplebenchcollected by Epoch AI0.5high
vpctcollected by Epoch AI0.5medium
weirdmlcollected by Epoch AI0.5high

These scores are not comparable between houses and are not added up. A score is dated by when it was evaluated, not when the model was released, and a score without its scaffold is not a measurement — which is why both are printed.

What retires, and when

  • deprecation2026-12-11, replaced by gpt-5.6-sol

Every rate we hold · Why there is no single best model ·