Speed leaderboard
The fastest LLMs, ranked
Ranked by median output tokens per second on OpenRouter. Speed is the difference between an app that feels instant and one that feels sluggish — critical for chat UIs, code completions and streaming responses.
Category leader
#01MiniMax
MiniMax: MiniMax M3.
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Speed
118
tok/s
Ranking
N=20| # | Model | Provider | Speed | Ctx | Released |
|---|---|---|---|---|---|
| 01 | MiniMax: MiniMax M3 | MiniMax | 118 | 1.0M | 2026-05-31 |
| 02 | xAI: Grok 4.3 | xAI | 109 | 1.0M | 2026-04-30 |
| 03 | OpenAI: GPT-5.6 Terra | OpenAI | 102 | 1.1M | 2026-07-09 |
| 04 | MoonshotAI: Kimi K2.6 | Moonshot | 98 | 262K | 2026-04-20 |
| 05 | NVIDIA: Nemotron 3 Ultra | NVIDIA | 96 | 1.0M | 2026-06-04 |
| 06 | MoonshotAI: Kimi K2.7 Code | Moonshot | 93 | 262K | 2026-06-12 |
| 07 | OpenAI: GPT-5.6 Sol | OpenAI | 92 | 1.1M | 2026-07-09 |
| 08 | DeepSeek: DeepSeek V4 Pro | DeepSeek | 91 | 1.0M | 2026-04-24 |
| 09 | OpenAI: GPT-5.6 Terra Pro | OpenAI | 89 | 1.1M | 2026-07-09 |
| 10 | Mistral: Mistral Medium 3.5 | Mistral | 86 | 262K | 2026-04-30 |
| 11 | Meta: Muse Spark 1.1 | Meta | 84 | 1.0M | 2026-07-16 |
| 12 | Anthropic: Claude Opus 4.8 (Fast) | Anthropic | 82 | 1.0M | 2026-05-27 |
| 13 | Z.ai: GLM 5.1 | Z.ai | 78 | 203K | 2026-04-07 |
| 14 | xAI: Grok 4.5 | xAI | 76 | 500K | 2026-07-08 |
| 15 | Z.ai: GLM 5.2 | Z.ai | 74 | 1.0M | 2026-06-16 |
| 16 | Qwen: Qwen3.7 Max | Alibaba | 68 | 1.0M | 2026-05-21 |
| 17 | MoonshotAI: Kimi K3 | Moonshot | 55 | 1.0M | 2026-07-16 |
| 18 | Anthropic: Claude Fable 5 | Anthropic | 42 | 1.0M | 2026-06-09 |
| 19 | Anthropic: Claude Opus 4.8 | Anthropic | 38 | 1.0M | 2026-05-27 |
| 20 | Sakana: Fugu Ultra | Sakana | 31 | 1.0M | 2026-06-24 |
Frequently asked
- Which LLM is the fastest?
- The top of this table shows the current leader in median output tokens per second. Speed matters most for interactive UIs, streaming responses and long generations where latency compounds.
- Is the fastest model the best one to use?
- Rarely — flagship reasoning models are usually slower than lite/flash variants. Pick fast models for high-throughput or latency-sensitive workloads; pick the SOTA when quality matters more than speed.