Price leaderboard
The cheapest frontier LLMs
Ranked by output price per million tokens. For high-volume workloads — classification, extraction, batch inference — the right frontier model can be 10× cheaper than the SOTA at similar quality.
Category leader
DeepSeek: DeepSeek V4 Flash 0731.
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.
Ranking
N=55| # | Model | Provider | Output price | Ctx | Released |
|---|---|---|---|---|---|
| 01 | DeepSeek: DeepSeek V4 Flash 0731 | DeepSeek | $0.18 | 1.0M | 2026-07-31 |
| 02 | DeepSeek: DeepSeek V4 Flash | DeepSeek | $0.28 | 1.0M | 2026-04-24 |
| 03 | Tencent: Hy3 | Tencent | $0.53 | 262K | 2026-07-06 |
| 04 | Meta: Muse Glimmer 30B | Meta | $1.10 | 131K | 2026-08-09 |
| 05 | StepFun: Step 3.7 Flash | StepFun | $1.15 | 262K | 2026-05-28 |
| 06 | MiniMax: MiniMax M3 | MiniMax | $1.20 | 1.0M | 2026-05-31 |
| 07 | Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | $1.50 | 66K | 2026-06-30 | |
| 08 | Qwen 3 Max | Alibaba | $1.50 | 262K | 2026-06-15 |
| 09 | Mistral: Mistral Large 3 2512 | Mistral | $1.50 | 262K | 2025-12-01 |
| 10 | Google: Gemini 3.1 Flash Lite | $1.50 | 1.0M | 2026-05-07 | |
| 11 | Qwen: Qwen3.5 Plus 2026-04-20 | Alibaba | $1.80 | 1.0M | 2026-04-27 |
| 12 | DeepSeek: DeepSeek V4 Pro 0423 | DeepSeek | $1.86 | 1.0M | 2026-04-24 |
| 13 | Z.ai: GLM 5 | Z.ai | $1.92 | 205K | 2026-02-11 |
| 14 | Qwen: Qwen3.6 Plus | Alibaba | $1.95 | 1.0M | 2026-04-02 |
| 15 | SpaceXAI: Grok Build 0.1 | xAI | $2.00 | 256K | 2026-05-20 |
| 16 | SpaceXAI: Grok 4.20 | xAI | $2.50 | 2.0M | 2026-03-31 |
| 17 | Llama 4 405B | Meta | $2.70 | 128K | 2026-04-30 |
| 18 | Google: Nano Banana 2 (Gemini 3.1 Flash Image) | $3.00 | 131K | 2026-06-18 | |
| 19 | Z.ai: GLM 5.2 | Z.ai | $3.04 | 1.0M | 2026-06-16 |
| 20 | DeepSeek: DeepSeek V4 Pro 0813 | DeepSeek | $3.36 | 1.0M | 2026-08-12 |
| 21 | Qwen: Qwen3.5 397B A17B | Alibaba | $3.60 | 262K | 2026-02-16 |
| 22 | Google: Gemini 3.8 Flash | $3.75 | 1.0M | 2026-09-02 | |
| 23 | Qwen: Qwen3 Max Thinking | Alibaba | $3.90 | 262K | 2026-02-09 |
| 24 | MoonshotAI: Kimi K2.6 | Moonshot | $4.00 | 262K | 2026-04-20 |
| 25 | Z.ai: GLM 5V Turbo | Z.ai | $4.00 | 203K | 2026-04-01 |
| 26 | Meta: Muse Spark 1.3 | Meta | $4.25 | 1.0M | 2026-09-02 |
| 27 | Meta: Muse Spark 1.1 | Meta | $4.25 | 1.0M | 2026-07-16 |
| 28 | Z.ai: GLM 5.3 | Z.ai | $4.40 | 1.3M | 2026-08-18 |
| 29 | Qwen: Qwen3.7 Max | Alibaba | $4.42 | 1.0M | 2026-05-21 |
| 30 | GPT-5.6 Terra | OpenAI | $4.80 | 400K | 2026-07-08 |
| 31 | Qwen: Qwen3.8 Max (0902) | Alibaba | $6.00 | 1.0M | 2026-09-03 |
| 32 | SpaceXAI: Grok 4.6 | xAI | $6.00 | 500K | 2026-08-12 |
| 33 | Qwen: Qwen3.8 2.4T A95B | Alibaba | $6.00 | 1.0M | 2026-08-12 |
| 34 | SpaceXAI: Grok 4.5 | xAI | $6.00 | 500K | 2026-07-08 |
| 35 | Mistral Large 3 | Mistral | $6.00 | 256K | 2026-05-08 |
| 36 | Qwen: Qwen3.6 Max Preview | Alibaba | $6.24 | 262K | 2026-04-27 |
| 37 | Mistral: Mistral Medium 3.5 | Mistral | $7.50 | 262K | 2026-04-30 |
| 38 | Grok 4 | xAI | $8.00 | 256K | 2026-06-11 |
| 39 | OpenAI: GPT-5.6 Sol | OpenAI | $10.00 | 1.1M | 2026-07-09 |
| 40 | Anthropic: Claude Sonnet 5 | Anthropic | $10.00 | 1.0M | 2026-06-30 |
| 41 | OpenAI: GPT-5.6 Sol Pro | OpenAI | $10.00 | 1.1M | 2026-07-09 |
| 42 | Google: Nano Banana Pro (Gemini 3 Pro Image) | $12.00 | 66K | 2026-06-18 | |
| 43 | OpenAI: GPT-5.6 Terra Pro | OpenAI | $12.00 | 1.1M | 2026-07-09 |
| 44 | MoonshotAI: Kimi K3 | Moonshot | $15.00 | 1.0M | 2026-07-16 |
| 45 | OpenAI: GPT-5.4 Image 2 | OpenAI | $15.00 | 272K | 2026-04-21 |
| 46 | OpenAI: GPT-5.4 | OpenAI | $15.00 | 1.1M | 2026-03-05 |
| 47 | Claude 4.5 Opus | Anthropic | $24.00 | 500K | 2026-05-19 |
| 48 | Claude Opus 5 | Anthropic | $25.00 | 1.0M | 2026-07-24 |
| 49 | Anthropic: Claude Opus 4.8 | Anthropic | $25.00 | 1.0M | 2026-05-27 |
| 50 | OpenAI: GPT-5.5 | OpenAI | $30.00 | 1.1M | 2026-04-24 |
| 51 | Sakana: Fugu Ultra | Sakana | $30.00 | 1.0M | 2026-06-24 |
| 52 | Anthropic: Claude Fable 5.1 | Anthropic | $50.00 | 1.0M | 2026-09-01 |
| 53 | Anthropic: Claude Fable 5 | Anthropic | $50.00 | 1.0M | 2026-06-09 |
| 54 | Anthropic: Claude Opus 4.8 (Fast) | Anthropic | $50.00 | 1.0M | 2026-05-27 |
| 55 | OpenAI: GPT-5.5 Pro | OpenAI | $180.00 | 1.1M | 2026-04-24 |
Why the cheapest model is often the wrong one
Per-token price is the most quoted and least useful number in this table. What you actually pay is price multiplied by tokens consumed, and modern reasoning models can emit five to twenty times as many tokens as a direct-answer model for the same visible output, because the hidden chain of thought is billed too. A model listed at a third of the price can easily cost more per completed task.
The second distortion is retries. A cheap model that gets a classification right 88% of the time, in a pipeline where an incorrect answer triggers a second attempt or human review, is competing against an expensive model at 97% with the full cost of that review loop attached. Work out the cost of a wrong answer before optimising the cost of a token.
Input and output are also priced differently, usually with output several times more expensive. Retrieval-heavy workloads that stuff long documents into context are dominated by input price, while generation-heavy work is dominated by output. Prompt caching changes the arithmetic again: if your system prompt is long and stable, a provider with aggressive caching can beat a nominally cheaper competitor outright.
Where cheap models genuinely win is high-volume, well-specified, verifiable work — classification, extraction, routing, summarising, tagging, first-pass triage — especially when a validator or schema catches failures automatically. The standard pattern is a cheap model doing the volume with a frontier model handling escalations, which typically lands within a few percent of frontier quality at a fraction of the spend.
Prices here are pulled from live provider listings and change without notice, so treat them as a shortlist tool rather than a quote. Run your real workload against two or three candidates, measure cost per successful outcome, and pick from that.
More on how these numbers are produced in the methodology and what each evaluation measures in the benchmark guide. Spotted a score that disagrees with its source? Tell us.
Frequently asked
- Which is the cheapest frontier LLM?
- The top of this table shows the model with the lowest output price per 1M tokens, among tracked frontier models. Cheaper models are ideal for high-volume classification, extraction, and batch workloads.
- Do cheaper models cost less overall?
- Not always — reasoning-heavy models generate more tokens per prompt (chain-of-thought, tool calls). Compare cost per completed task, not just per-token price.