Long-context leaderboard
LLMs with the largest context windows
Ranked by maximum input tokens the model accepts in a single call. Larger windows unlock whole-codebase analysis, long-document summarization and multi-hour transcripts — though effective recall often lags the advertised window.
Category leader
SpaceXAI: Grok 4.20.
Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...
Ranking
N=55| # | Model | Provider | Context | Ctx | Released |
|---|---|---|---|---|---|
| 01 | SpaceXAI: Grok 4.20 | xAI | 2.0M | 2.0M | 2026-03-31 |
| 02 | Z.ai: GLM 5.3 | Z.ai | 1.3M | 1.3M | 2026-08-18 |
| 03 | OpenAI: GPT-5.6 Sol | OpenAI | 1.1M | 1.1M | 2026-07-09 |
| 04 | OpenAI: GPT-5.5 Pro | OpenAI | 1.1M | 1.1M | 2026-04-24 |
| 05 | OpenAI: GPT-5.5 | OpenAI | 1.1M | 1.1M | 2026-04-24 |
| 06 | OpenAI: GPT-5.4 | OpenAI | 1.1M | 1.1M | 2026-03-05 |
| 07 | OpenAI: GPT-5.6 Terra Pro | OpenAI | 1.1M | 1.1M | 2026-07-09 |
| 08 | OpenAI: GPT-5.6 Sol Pro | OpenAI | 1.1M | 1.1M | 2026-07-09 |
| 09 | MoonshotAI: Kimi K3 | Moonshot | 1.0M | 1.0M | 2026-07-16 |
| 10 | Meta: Muse Spark 1.3 | Meta | 1.0M | 1.0M | 2026-09-02 |
| 11 | DeepSeek: DeepSeek V4 Pro 0813 | DeepSeek | 1.0M | 1.0M | 2026-08-12 |
| 12 | Qwen: Qwen3.8 2.4T A95B | Alibaba | 1.0M | 1.0M | 2026-08-12 |
| 13 | Z.ai: GLM 5.2 | Z.ai | 1.0M | 1.0M | 2026-06-16 |
| 14 | DeepSeek: DeepSeek V4 Pro 0423 | DeepSeek | 1.0M | 1.0M | 2026-04-24 |
| 15 | Meta: Muse Spark 1.1 | Meta | 1.0M | 1.0M | 2026-07-16 |
| 16 | Google: Gemini 3.8 Flash | 1.0M | 1.0M | 2026-09-02 | |
| 17 | DeepSeek: DeepSeek V4 Flash 0731 | DeepSeek | 1.0M | 1.0M | 2026-07-31 |
| 18 | DeepSeek: DeepSeek V4 Flash | DeepSeek | 1.0M | 1.0M | 2026-04-24 |
| 19 | MiniMax: MiniMax M3 | MiniMax | 1.0M | 1.0M | 2026-05-31 |
| 20 | Google: Gemini 3.1 Flash Lite | 1.0M | 1.0M | 2026-05-07 | |
| 21 | Anthropic: Claude Fable 5.1 | Anthropic | 1.0M | 1.0M | 2026-09-01 |
| 22 | Qwen: Qwen3.8 Max (0902) | Alibaba | 1.0M | 1.0M | 2026-09-03 |
| 23 | Claude Opus 5 | Anthropic | 1.0M | 1.0M | 2026-07-24 |
| 24 | Anthropic: Claude Fable 5 | Anthropic | 1.0M | 1.0M | 2026-06-09 |
| 25 | Anthropic: Claude Opus 4.8 | Anthropic | 1.0M | 1.0M | 2026-05-27 |
| 26 | Qwen: Qwen3.7 Max | Alibaba | 1.0M | 1.0M | 2026-05-21 |
| 27 | Anthropic: Claude Sonnet 5 | Anthropic | 1.0M | 1.0M | 2026-06-30 |
| 28 | Qwen: Qwen3.6 Plus | Alibaba | 1.0M | 1.0M | 2026-04-02 |
| 29 | Anthropic: Claude Opus 4.8 (Fast) | Anthropic | 1.0M | 1.0M | 2026-05-27 |
| 30 | Qwen: Qwen3.5 Plus 2026-04-20 | Alibaba | 1.0M | 1.0M | 2026-04-27 |
| 31 | Sakana: Fugu Ultra | Sakana | 1.0M | 1.0M | 2026-06-24 |
| 32 | SpaceXAI: Grok 4.6 | xAI | 500K | 500K | 2026-08-12 |
| 33 | SpaceXAI: Grok 4.5 | xAI | 500K | 500K | 2026-07-08 |
| 34 | Claude 4.5 Opus | Anthropic | 500K | 500K | 2026-05-19 |
| 35 | GPT-5.6 Terra | OpenAI | 400K | 400K | 2026-07-08 |
| 36 | OpenAI: GPT-5.4 Image 2 | OpenAI | 272K | 272K | 2026-04-21 |
| 37 | MoonshotAI: Kimi K2.6 | Moonshot | 262K | 262K | 2026-04-20 |
| 38 | Qwen: Qwen3 Max Thinking | Alibaba | 262K | 262K | 2026-02-09 |
| 39 | Mistral: Mistral Medium 3.5 | Mistral | 262K | 262K | 2026-04-30 |
| 40 | Tencent: Hy3 | Tencent | 262K | 262K | 2026-07-06 |
| 41 | StepFun: Step 3.7 Flash | StepFun | 262K | 262K | 2026-05-28 |
| 42 | Qwen: Qwen3.5 397B A17B | Alibaba | 262K | 262K | 2026-02-16 |
| 43 | Qwen 3 Max | Alibaba | 262K | 262K | 2026-06-15 |
| 44 | Mistral: Mistral Large 3 2512 | Mistral | 262K | 262K | 2025-12-01 |
| 45 | Qwen: Qwen3.6 Max Preview | Alibaba | 262K | 262K | 2026-04-27 |
| 46 | Grok 4 | xAI | 256K | 256K | 2026-06-11 |
| 47 | SpaceXAI: Grok Build 0.1 | xAI | 256K | 256K | 2026-05-20 |
| 48 | Mistral Large 3 | Mistral | 256K | 256K | 2026-05-08 |
| 49 | Z.ai: GLM 5 | Z.ai | 205K | 205K | 2026-02-11 |
| 50 | Z.ai: GLM 5V Turbo | Z.ai | 203K | 203K | 2026-04-01 |
| 51 | Meta: Muse Glimmer 30B | Meta | 131K | 131K | 2026-08-09 |
| 52 | Google: Nano Banana 2 (Gemini 3.1 Flash Image) | 131K | 131K | 2026-06-18 | |
| 53 | Llama 4 405B | Meta | 128K | 128K | 2026-04-30 |
| 54 | Google: Nano Banana Pro (Gemini 3 Pro Image) | 66K | 66K | 2026-06-18 | |
| 55 | Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | 66K | 66K | 2026-06-30 |
Advertised context is not usable context
A context window is a ceiling, not a promise. Every model degrades before it reaches its advertised limit, and the shape of that degradation matters more than the headline number. The well-documented pattern is that material at the very start and very end of a long prompt is recalled reliably while material in the middle is often missed — the lost-in-the-middle effect — which means a million-token window does not mean a million tokens of dependable attention.
Retrieval benchmarks flatter models here. Finding one planted sentence in a long document is far easier than reasoning across twelve facts scattered through it, and most published long-context results measure the former. If your task needs synthesis across a whole corpus rather than lookup within it, expect real performance well below the advertised capability.
Long prompts are also expensive and slow. Input tokens are billed on every call, so a large stable prefix multiplies cost across a conversation unless the provider caches it, and time to first token grows with prompt length. A retrieval step that sends four relevant pages instead of four hundred is usually faster, cheaper and more accurate than relying on the window.
Where large windows genuinely earn their place: whole-repository code understanding, long legal or financial documents that resist clean chunking, multi-hour transcripts, and agent loops that accumulate long tool-call histories. In those cases the window removes an entire class of chunking bugs, and that is worth paying for.
Practical rule — use retrieval by default, use the long window when chunking would destroy the structure you need, and always test recall on your own documents at the length you actually intend to use, not at the length the specification allows.
More on how these numbers are produced in the methodology and what each evaluation measures in the benchmark guide. Spotted a score that disagrees with its source? Tell us.
Frequently asked
- Which LLM has the largest context window?
- The current leader is at the top of this table. Context windows range from 128K to 2M+ tokens across frontier models; larger windows unlock whole-codebase and full-document workflows.
- Is bigger context always better?
- Not necessarily. Recall quality often degrades past ~200K tokens even when the window advertises 1M+. Use benchmarks like Needle-in-a-Haystack to verify effective context, not just declared window size.