OpenAI

gpt-5.4-2026-03-05

Made by OpenAI. Released 2026-03-05. Accessibility not recorded. We hold no price for this model: no seller we can read lists it any more, so what follows is what was measured, not what it cost.

What it scores, by house

BenchmarkHouseScoreEvaluatedScaffold
chess puzzlesEpoch AI0.2±0.042026-07-15low
chess puzzlesEpoch AI0.4±0.052026-07-15high
chess puzzlesEpoch AI0.1±0.022026-07-15none
chess puzzlesEpoch AI0.4±0.052026-07-15medium
chess puzzlesEpoch AI0.4±0.052026-03-11xhigh
ebr benchEpoch AI0.32026-06-25xhigh
frontiermathEpoch AI0.5±0.032026-03-06xhigh
frontiermath tier 4Epoch AI0.3±0.062026-03-06xhigh
frontiermath tier 4 v2Epoch AI0.5±0.082026-06-11xhigh
frontiermath tiers 1 3 v2Epoch AI0.8±0.022026-06-11xhigh
gpqa diamondEpoch AI0.8±0.032026-07-15low
gpqa diamondEpoch AI0.7±0.032026-07-15none
gpqa diamondEpoch AI0.9±0.022026-07-15medium
gpqa diamondEpoch AI0.9±0.022026-07-15high
gpqa diamondEpoch AI0.9±0.022026-03-06xhigh
mirrorcodeEpoch AI0.2±0.082026-08-10high
mystery game puzzlesEpoch AI0.2±0.042026-08-30none
mystery game puzzlesEpoch AI0.3±0.052026-08-28medium
mystery game puzzlesEpoch AI0.2±0.042026-08-27low
mystery game puzzlesEpoch AI0.4±0.052026-07-24xhigh
otis mock aime 2024 2025near its ceilingEpoch AI1.0±0.032026-07-15medium
otis mock aime 2024 2025near its ceilingEpoch AI0.6±0.072026-07-15none
otis mock aime 2024 2025near its ceilingEpoch AI1.0±0.022026-07-15high
otis mock aime 2024 2025near its ceilingEpoch AI0.8±0.052026-07-15low
otis mock aime 2024 2025near its ceilingEpoch AI1.0±0.032026-03-06xhigh
simpleqa verifiedEpoch AI0.5±0.022026-08-27xhigh
swe bench verifiedEpoch AI0.8±0.022026-03-06high
ale benchcollected by Epoch AI1086none
ale benchcollected by Epoch AI1607high
ale benchcollected by Epoch AI1521medium
algotunecollected by Epoch AI1.9high
apex agentscollected by Epoch AI0.5unknown
apex agentscollected by Epoch AI0.5xhigh
arc aginear its ceilingcollected by Epoch AI0.9medium
arc aginear its ceilingcollected by Epoch AI0.9high
arc aginear its ceilingcollected by Epoch AI0.9xhigh
arc aginear its ceilingcollected by Epoch AI0.7low
arc agi 2collected by Epoch AI0.7high
arc agi 2collected by Epoch AI0.6medium
arc agi 2collected by Epoch AI0.3low
arc agi 2collected by Epoch AI0.7xhigh
blueprint bench 2collected by Epoch AI0.3unknown
cl benchcollected by Epoch AI0.3xhigh
cl bench lifecollected by Epoch AI0.1unknown
cl bench lifecollected by Epoch AI0.2high
cl bench lifecollected by Epoch AI0.2xhigh
critptcollected by Epoch AI0.2xhigh
deepresearchbenchcollected by Epoch AI0.4low
deepswecollected by Epoch AI0.5mini-swe-agent
dtbenchnear its ceilingcollected by Epoch AI0.9xhigh
enigma evalcollected by Epoch AI0.2xhigh
forecastbenchcollected by Epoch AI59.5unknown
gbaevalcollected by Epoch AI0.5unknown
gsocollected by Epoch AI0.3OpenHands
hlecollected by Epoch AI0.4xhigh
lmcacollected by Epoch AI52.0xhigh
metr time horizonscollected by Epoch AI342xhigh
metr time horizonscollected by Epoch AI342unknown
posttrainbenchcollected by Epoch AI0.2Codex CLI
proofbenchcollected by Epoch AI0.6xhigh
scicodecollected by Epoch AI0.6xhigh
terminalbenchcollected by Epoch AI0.8ForgeCode
vending bench 2collected by Epoch AI6144unknown
webdev arenacollected by Epoch AI1391unknown
webdev arenacollected by Epoch AI1463high
webdev arenacollected by Epoch AI1442medium
weirdmlcollected by Epoch AI0.8xhigh
weirdmlcollected by Epoch AI0.6none

These scores are not comparable between houses and are not added up. A score is dated by when it was evaluated, not when the model was released, and a score without its scaffold is not a measurement — which is why both are printed.

Every rate we hold · Why there is no single best model ·