Speed leaderboard
The fastest LLMs, ranked
Ranked by median output tokens per second on OpenRouter. Speed is the difference between an app that feels instant and one that feels sluggish — critical for chat UIs, code completions and streaming responses.
Category leader
Google: Gemini 3.8 Flash.
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Ranking
N=48| # | Model | Provider | Speed | Ctx | Released |
|---|---|---|---|---|---|
| 01 | Google: Gemini 3.8 Flash | 190 | 1.0M | 2026-09-02 | |
| 02 | DeepSeek: DeepSeek V4 Flash 0731 | DeepSeek | 154 | 1.0M | 2026-07-31 |
| 03 | DeepSeek: DeepSeek V4 Flash | DeepSeek | 145 | 1.0M | 2026-04-24 |
| 04 | SpaceXAI: Grok 4.20 | xAI | 132 | 2.0M | 2026-03-31 |
| 05 | Qwen: Qwen3.5 Plus 2026-04-20 | Alibaba | 129 | 1.0M | 2026-04-27 |
| 06 | SpaceXAI: Grok Build 0.1 | xAI | 125 | 256K | 2026-05-20 |
| 07 | SpaceXAI: Grok 4.5 | xAI | 121 | 500K | 2026-07-08 |
| 08 | GPT-5.6 Terra | OpenAI | 118 | 400K | 2026-07-08 |
| 09 | SpaceXAI: Grok 4.6 | xAI | 116 | 500K | 2026-08-12 |
| 10 | OpenAI: GPT-5.6 Sol | OpenAI | 109 | 1.1M | 2026-07-09 |
| 11 | Qwen 3 Max | Alibaba | 105 | 262K | 2026-06-15 |
| 12 | Meta: Muse Glimmer 30B | Meta | 103 | 131K | 2026-08-09 |
| 13 | MiniMax: MiniMax M3 | MiniMax | 96 | 1.0M | 2026-05-31 |
| 14 | Z.ai: GLM 5V Turbo | Z.ai | 96 | 203K | 2026-04-01 |
| 15 | Anthropic: Claude Sonnet 5 | Anthropic | 95 | 1.0M | 2026-06-30 |
| 16 | Grok 4 | xAI | 95 | 256K | 2026-06-11 |
| 17 | OpenAI: GPT-5.5 | OpenAI | 91 | 1.1M | 2026-04-24 |
| 18 | Mistral Large 3 | Mistral | 88 | 256K | 2026-05-08 |
| 19 | DeepSeek: DeepSeek V4 Pro 0423 | DeepSeek | 86 | 1.0M | 2026-04-24 |
| 20 | Z.ai: GLM 5 | Z.ai | 84 | 205K | 2026-02-11 |
| 21 | Qwen: Qwen3.7 Max | Alibaba | 83 | 1.0M | 2026-05-21 |
| 22 | Tencent: Hy3 | Tencent | 83 | 262K | 2026-07-06 |
| 23 | Anthropic: Claude Opus 4.8 (Fast) | Anthropic | 83 | 1.0M | 2026-05-27 |
| 24 | DeepSeek: DeepSeek V4 Pro 0813 | DeepSeek | 82 | 1.0M | 2026-08-12 |
| 25 | Mistral: Mistral Medium 3.5 | Mistral | 82 | 262K | 2026-04-30 |
| 26 | Qwen: Qwen3.8 2.4T A95B | Alibaba | 77 | 1.0M | 2026-08-12 |
| 27 | Z.ai: GLM 5.2 | Z.ai | 76 | 1.0M | 2026-06-16 |
| 28 | Llama 4 405B | Meta | 74 | 128K | 2026-04-30 |
| 29 | Qwen: Qwen3.8 Max (0902) | Alibaba | 73 | 1.0M | 2026-09-03 |
| 30 | Z.ai: GLM 5.3 | Z.ai | 71 | 1.3M | 2026-08-18 |
| 31 | Qwen: Qwen3 Max Thinking | Alibaba | 71 | 262K | 2026-02-09 |
| 32 | Claude Opus 5 | Anthropic | 68 | 1.0M | 2026-07-24 |
| 33 | Anthropic: Claude Opus 4.8 | Anthropic | 65 | 1.0M | 2026-05-27 |
| 34 | Meta: Muse Spark 1.3 | Meta | 64 | 1.0M | 2026-09-02 |
| 35 | Qwen: Qwen3.5 397B A17B | Alibaba | 64 | 262K | 2026-02-16 |
| 36 | Mistral: Mistral Large 3 2512 | Mistral | 64 | 262K | 2025-12-01 |
| 37 | Claude 4.5 Opus | Anthropic | 64 | 500K | 2026-05-19 |
| 38 | Anthropic: Claude Fable 5.1 | Anthropic | 63 | 1.0M | 2026-09-01 |
| 39 | MoonshotAI: Kimi K2.6 | Moonshot | 63 | 262K | 2026-04-20 |
| 40 | Qwen: Qwen3.6 Max Preview | Alibaba | 62 | 262K | 2026-04-27 |
| 41 | Anthropic: Claude Fable 5 | Anthropic | 61 | 1.0M | 2026-06-09 |
| 42 | Meta: Muse Spark 1.1 | Meta | 61 | 1.0M | 2026-07-16 |
| 43 | OpenAI: GPT-5.4 | OpenAI | 59 | 1.1M | 2026-03-05 |
| 44 | MoonshotAI: Kimi K3 | Moonshot | 58 | 1.0M | 2026-07-16 |
| 45 | OpenAI: GPT-5.6 Terra Pro | OpenAI | 52 | 1.1M | 2026-07-09 |
| 46 | OpenAI: GPT-5.5 Pro | OpenAI | 42 | 1.1M | 2026-04-24 |
| 47 | Sakana: Fugu Ultra | Sakana | 39 | 1.0M | 2026-06-24 |
| 48 | OpenAI: GPT-5.6 Sol Pro | OpenAI | 37 | 1.1M | 2026-07-09 |
What throughput does and does not tell you
Speed has two components that behave very differently. Time to first token is the pause before anything appears, and it is what users experience as lag. Tokens per second is how fast text streams once it starts. A model with excellent throughput and a slow cold start still feels sluggish in a chat interface, while a model that starts instantly can feel fast even at modest streaming rates.
Reasoning models complicate the picture further. Hidden thinking tokens are generated before the visible answer begins, so a model can post a high tokens-per-second figure while taking far longer than a slower-streaming competitor to actually deliver a response. If your product is interactive, measure wall-clock time to a complete useful answer, not the raw rate.
These figures are measurements against live endpoints, which means they move with provider, region, hardware, quantisation and time of day. The same model served by two providers can differ substantially, and a heavily loaded endpoint at peak hours will not match a benchmark run at 3am. Treat the ordering as reliable and the absolute numbers as indicative.
For interactive products, speed deserves as much weight as quality — an answer that arrives after the user has given up has failed regardless of how good it was. For batch pipelines, throughput barely matters compared with concurrency limits and cost per completed task, and it is usually a mistake to trade accuracy for it.
The pragmatic pattern is tiering: a fast model for the first response or the interactive path, and a slower frontier model behind it for the work that genuinely needs the extra reasoning. Most products that feel quick and smart at the same time are doing exactly this.
More on how these numbers are produced in the methodology and what each evaluation measures in the benchmark guide. Spotted a score that disagrees with its source? Tell us.
Frequently asked
- Which LLM is the fastest?
- The top of this table shows the current leader in median output tokens per second. Speed matters most for interactive UIs, streaming responses and long generations where latency compounds.
- Is the fastest model the best one to use?
- Rarely — flagship reasoning models are usually slower than lite/flash variants. Pick fast models for high-throughput or latency-sensitive workloads; pick the SOTA when quality matters more than speed.