Which models matter right now?

The current model in each major lab's main lines, one card each: what it costs, which independent tests it comes first on, where nobody has measured it yet, and the cheaper model that comes closest.

15flagship models from 8 labs, picked by rule: the newest version of each lab's main lines. 6 have not been scored on the Board's tests yet, and their cards say so.

Not a ranking: labs are in a fixed order and nothing here adds scores from different tests together. For who leads each test, see Which is best?; for what a job costs on each, see the cost page.

OpenAI

GPT-6 Astrafirst on 8 tests

OpenAI's dearest line · released 3 Sept 2026

$15.00an hour of agent work

$10 in / $50 out per million tokens

First on
  • GPQA Diamond95.8%of 213
  • Chess puzzles72.0%of 137
  • FrontierMath Tiers 1–3 (v2)93.7%of 75
  • SimpleQA75.6%of 71
  • Mystery games84.0%of 67

and 3 more on the Board

Top three
  • MirrorCode#346.7%of 8

measured on 10 of 14 tests · latest evaluated 30 Aug 2026

Nearly as good for lessGemini 3.8 Flash scores 95.4% to this model's 95.8% on GPQA Diamond, at $1.13 an hour — 13× cheaper.

GPT-6 Lunanot measured yet

OpenAI's cheapest line

$0.150an hour of agent work

$0.10 in / $0.50 out per million tokens

No independent scores under this name yet. Epoch AI has not published results for a model called GPT-6 Luna; the newest models usually reach its record a few weeks after they go on sale.

Anthropic

Claude Opus 5.5not measured yet

Anthropic's cheaper line

$6.00an hour of agent work

$4 in / $20 out per million tokens

No independent scores under this name yet. Epoch AI has not published results for a model called Claude Opus 5.5; the newest models usually reach its record a few weeks after they go on sale.

Claude Fable 5.1first on 1 test

Anthropic's dearest line · released 1 Sept 2026

$15.00an hour of agent work

$10 in / $50 out per million tokens

First on
  • MirrorCode73.3%of 8
Top three
  • FrontierMath Tiers 1–3 (v2)#290.2%of 75
  • EBR-Bench#257.1%of 21
  • SimpleQA#370.8%of 71
  • FrontierMath Tier 4 (v2)#387.8%of 57

measured on 9 of 14 tests · latest evaluated 10 Sept 2026

Google

Gemini 3.1 Protop three on 1

Google's dearest line · released 19 Feb 2026

$3.20an hour of agent work

$2 in / $12 out per million tokens

Top three
  • SimpleQA#273.5%of 71

measured on 11 of 14 tests · latest evaluated 10 Aug 2026

Gemini 3.8 Flashtop three on 1

Google's cheaper line · released 2 Sept 2026

$1.13an hour of agent work

$0.75 in / $3.75 out per million tokens

Top three
  • GPQA Diamond#295.4%of 213

measured on 7 of 14 tests · latest evaluated 2 Sept 2026

Nearly as good for lessDeepSeek V4 Pro scores 90.9% to this model's 95.4% on GPQA Diamond, at $0.858 an hour — 1.3× cheaper.

xAI

DeepSeek

DeepSeek V4 Promeasured on 8

DeepSeek's dearest line · released 24 Apr 2026

$0.858an hour of agent work

$0.66 in / $1.98 out per million tokens

Not in the top three on any of the Board's open tests.

measured on 8 of 14 tests · latest evaluated 27 Aug 2026

DeepSeek Flashnot measured yet

DeepSeek's cheaper line

$0.210an hour of agent work

$0.15 in / $0.60 out per million tokens

No independent scores under this name yet. Epoch AI has not published results for a model called DeepSeek Flash; the newest models usually reach its record a few weeks after they go on sale.

Moonshot

Kimi K3measured on 7

Moonshot's flagship line · released 16 Jul 2026

$4.50an hour of agent work

$3 in / $15 out per million tokens

Not in the top three on any of the Board's open tests.

measured on 7 of 14 tests · latest evaluated 29 Aug 2026

Nearly as good for lessGLM-5.3 Flash scores 90.2% to this model's 93.1% on GPQA Diamond, at $0.200 an hour — 23× cheaper.

Z.ai

GLM-5.3measured on 7

Z.ai's dearest line · released 14 Aug 2026

$1.84an hour of agent work

$1.40 in / $4.40 out per million tokens

Not in the top three on any of the Board's open tests.

measured on 7 of 14 tests · latest evaluated 30 Aug 2026

As good or better for lessGemini 3.8 Flash scores 47.0% to this model's 33.0% on Mystery games, at $1.13 an hour — 1.6× cheaper.

GLM-5.3 Flashmeasured on 6

Z.ai's cheaper line · released 20 Aug 2026

$0.200an hour of agent work

$0.15 in / $0.50 out per million tokens

Not in the top three on any of the Board's open tests.

measured on 6 of 14 tests · latest evaluated 28 Aug 2026

Alibaba

Qwen3.8 Maxmeasured on 7

Alibaba's dearest line · released 2 Aug 2026

$2.60an hour of agent work

$2 in / $6 out per million tokens

Not in the top three on any of the Board's open tests.

measured on 7 of 14 tests · latest evaluated 27 Aug 2026

As good or better for lessGemini 3.8 Flash scores 47.0% to this model's 38.0% on Mystery games, at $1.13 an hour — 2.3× cheaper.

Qwen3.8 Flashnot measured yet

Alibaba's cheaper line · released 26 Aug 2026

$0.197an hour of agent work

$0.15 in / $0.47 out per million tokens

Epoch AI holds this model as qwen3.8-flash but has not scored it on any of the Board's tests yet.

How this page is made

Which models.Each lab's main product lines are named once — GPT Astra, Sol and Luna; Claude Opus and Fable; Gemini Pro and Flash, and so on — and the newest version of each that the lab still sells is picked automatically from its own price page. Speed and high-compute tiers of a line are left out.

First on, top three on. These are the ranks on the Board: first among every model Epoch AI has scored on that test, each with its evaluation date. Tests most models have already solved are left out, because coming first on one says little.

Nearly as good for less.On the model's best-ranked test, the cheapest other flagship scoring at least 95% of its score that also costs less for an hour of agent work. One test, named, with both scores — never an average across tests.

An hour of agent work is a million tokens in and a hundred thousand out at list price, as on the cost page. The bar is on a log scale.

Benchmark data from Epoch AI, used under CC-BY 4.0. Headlines are Datum's own, matched to a model only when they name it.