OpenAI

gpt-5-mini-2025-08-07

Made by OpenAI. Released 2025-08-07. Accessibility not recorded. We hold no price for this model: no seller we can read lists it any more, so what follows is what was measured, not what it cost.

What it scores, by house

BenchmarkHouseScoreEvaluatedScaffold
chess puzzlesEpoch AI0.3±0.052026-08-07high
chess puzzlesEpoch AI0.1±0.032026-08-07low
chess puzzlesEpoch AI0.1±0.032026-08-07minimal
frontiermathEpoch AI0.3±0.032025-11-13high
frontiermathEpoch AI0.2±0.022025-11-13medium
frontiermath tier 4Epoch AI0.1±0.042025-10-30high
frontiermath tier 4Epoch AI0.0±0.032025-08-07medium
frontiermath tier 4 v2Epoch AI0.1±0.052026-06-12high
frontiermath tiers 1 3 v2Epoch AI0.2±0.022026-08-27low
frontiermath tiers 1 3 v2Epoch AI0.1±0.012026-08-27minimal
frontiermath tiers 1 3 v2Epoch AI0.5±0.032026-06-12high
gpqa diamondEpoch AI0.7±0.032026-08-07minimal
gpqa diamondEpoch AI0.8±0.022025-10-30high
gpqa diamondEpoch AI0.7±0.022025-08-07medium
math level 5near its ceilingEpoch AI1.0±0.002025-10-30high
math level 5near its ceilingEpoch AI1.0±0.002025-08-20medium
mystery game puzzlesEpoch AI0.0±0.022026-08-27medium
mystery game puzzlesEpoch AI0.1±0.022026-08-27high
mystery game puzzlesEpoch AI0.1±0.032026-08-27minimal
otis mock aime 2024 2025near its ceilingEpoch AI0.6±0.072026-08-07minimal
otis mock aime 2024 2025near its ceilingEpoch AI0.9±0.042025-10-30high
otis mock aime 2024 2025near its ceilingEpoch AI0.8±0.052025-08-07medium
simpleqa verifiedEpoch AI0.2±0.012026-08-10high
swe bench verifiedEpoch AI0.6±0.022026-02-01medium
ale benchcollected by Epoch AI800high
algotunecollected by Epoch AI1.4high
arc aginear its ceilingcollected by Epoch AI0.3low
arc aginear its ceilingcollected by Epoch AI0.4medium
arc aginear its ceilingcollected by Epoch AI0.1minimal
arc aginear its ceilingcollected by Epoch AI0.5high
arc agi 2collected by Epoch AI0.0minimal
arc agi 2collected by Epoch AI0.0medium
arc agi 2collected by Epoch AI0.0high
arc agi 2collected by Epoch AI0.0low
critptcollected by Epoch AI0.0unknown
dtbenchnear its ceilingcollected by Epoch AI0.8high
enigma evalcollected by Epoch AI0.1unknown
fictionlivebenchcollected by Epoch AI0.6medium
forecastbenchcollected by Epoch AI61.0unknown
hlecollected by Epoch AI0.2unknown
lech mazur writingcollected by Epoch AI8.3medium
lmcacollected by Epoch AI34.2high
proofbenchcollected by Epoch AI0.1high
scicodecollected by Epoch AI0.4unknown
terminalbenchcollected by Epoch AI0.2Mini-SWE-Agent
terminalbenchcollected by Epoch AI0.3spoox-o-m
terminalbenchcollected by Epoch AI0.3spoox-m
terminalbenchcollected by Epoch AI0.3Codex CLI
terminalbenchcollected by Epoch AI0.3OpenHands
terminalbenchcollected by Epoch AI0.2Terminus 2
vending bench 2collected by Epoch AI-31.2unknown
vpctcollected by Epoch AI0.4medium
vpctcollected by Epoch AI0.4high
weirdmlcollected by Epoch AI0.5high

These scores are not comparable between houses and are not added up. A score is dated by when it was evaluated, not when the model was released, and a score without its scaffold is not a measurement — which is why both are printed.

What retires, and when

  • deprecation2026-12-11, replaced by gpt-5.6-terra

Every rate we hold · Why there is no single best model ·