xAI
grok-4.20-0309-reasoning
Made by xAI. Released 2026-02-17. Accessibility not recorded. Cheapest standard rate we hold: $2.50 per million output tokens, at < 200k prompt tokens context.
What it costs
| Tier | Kind | In | Out | Cached in | From |
|---|---|---|---|---|---|
| standard | text · < 200k prompt tokens | $1.25 | $2.50 | $0.20 | seller |
| standard | text · ≥ 200k prompt tokens | $2.50 | $5 | $0.40 | seller |
per million tokens, US dollars. * marks a rate from a single registry or two that disagree — the best available number rather than a confirmed one. Every other row is confirmed against a seller or a licensed index.
What it scores, by house
| Benchmark | House | Score | Evaluated | Scaffold |
|---|---|---|---|---|
| chess puzzles | Epoch AI | 0.2±0.04 | 2026-07-13 | — |
| frontiermath tier 4 v2 | Epoch AI | 0.2±0.06 | 2026-07-13 | — |
| frontiermath tiers 1 3 v2 | Epoch AI | 0.4±0.03 | 2026-07-13 | — |
| gpqa diamond | Epoch AI | 0.9±0.02 | 2026-07-13 | — |
| otis mock aime 2024 2025near its ceiling | Epoch AI | 0.9±0.03 | 2026-07-13 | — |
| simpleqa verified | Epoch AI | 0.3±0.01 | 2026-08-27 | — |
| forecastbench | collected by Epoch AI | 60.7 | — | — |
| proofbench | collected by Epoch AI | 0.1 | — | — |
These scores are not comparable between houses and are not added up. A score is dated by when it was evaluated, not when the model was released, and a score without its scaffold is not a measurement — which is why both are printed.
Every rate we hold · Why there is no single best model · Everything from xAI