What would it cost to actually do something?
Prices are quoted per million tokens, which is how the industry sells them and almost impossible to feel. So here is the same arithmetic in the unit a person has: one question, one three-hundred-page document, one day of an agent working. Pick a job and the flagship models are ranked by what it would cost you.
many turns, files read over and over
| Model | Seller | Run a coding agent for an hour | Once a day, for a year | Rate |
|---|---|---|---|---|
| GLM-5.3 Flash | Z.ai | $0.200 | $73.00 | $0.15 in / $0.50 out |
| DeepSeek Flash | DeepSeek | $0.210 | $76.65 | $0.15 in / $0.60 out |
| Gemini 3.8 Flash | $1.13 | $411 | $0.75 in / $3.75 out | |
| Kimi K2.6 | Moonshot | $1.35 | $493 | $0.95 in / $4 out |
| GLM-5.3 | Z.ai | $1.84 | $672 | $1.40 in / $4.40 out |
| Grok 4.6 | xAI | $2.60 | $949 | $2 in / $6 out |
| Gemini 3.1 Pro | $3.20 | $1168 | $2 in / $12 out | |
| Kimi K3 | Moonshot | $4.50 | $1643 | $3 in / $15 out |
| GPT-5.6 Sol | OpenAI | $6.00 | $2190 | $4 in / $20 out |
| Claude Opus 5 | Anthropic | $7.50 | $2738 | $5 in / $25 out |
| Claude Fable 5.1 | Anthropic | $15.00 | $5475 | $10 in / $50 out |
| GPT-6 Astra | OpenAI | $15.00 | $5475 | $10 in / $50 out |
1,000,000 tokens in, 100,000 out — an agent re-reads its context on every turn, so input dominates. A real agent run is usually much cheaper than this because most of that input is served from cache, which this does not count. These are published list rates only — no cached input, no batch discount, no retries and no tool calls — so treat every figure as an upper bound. Token counts differ between models, and a model's own rate can change: the table above says which prices are confirmed and which are not.
What this is not
It is the price of the tokens, not the price of the job. Nobody bills a question at list rate: caching, batch discounts, retries and the four rejected attempts before a good answer are all real, and none of them are in these numbers. Read them as the floor — the part of the bill that scales with what you send and what comes back.
And it says nothing about whether the output is any good. This site holds prices and it holds capability scores, and it keeps them apart on purpose: a cheap model that cannot do the job is not cheaper. The Board is where to look for which model is actually better, house by house, and the Gap for how far the downloadable ones are behind.
Looking for a rate rather than a job? Every published price we hold, filterable by seller, is on the rates page.