Release timeline
The newest LLMs, freshest first
Ranked by release date. New flagship models drop every few weeks — this list is the fastest way to see what just shipped from OpenAI, Anthropic, Google, xAI, DeepSeek, Meta, Kimi and others.
Category leader
Qwen: Qwen3.8 Max (0902).
Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...
Ranking
N=59| # | Model | Provider | Released | Ctx | Released |
|---|---|---|---|---|---|
| 01 | Qwen: Qwen3.8 Max (0902) | Alibaba | 2026-09-03 | 1.0M | 2026-09-03 |
| 02 | Meta: Muse Spark 1.3 | Meta | 2026-09-02 | 1.0M | 2026-09-02 |
| 03 | Google: Gemini 3.8 Flash | 2026-09-02 | 1.0M | 2026-09-02 | |
| 04 | Anthropic: Claude Fable 5.1 | Anthropic | 2026-09-01 | 1.0M | 2026-09-01 |
| 05 | Tencent: Hy4 preview | Tencent | 2026-08-28 | 1.0M | 2026-08-28 |
| 06 | Z.ai: GLM 5.3 | Z.ai | 2026-08-18 | 1.3M | 2026-08-18 |
| 07 | Qwen: Qwen3.8 2.4T A95B | Alibaba | 2026-08-12 | 1.0M | 2026-08-12 |
| 08 | DeepSeek: DeepSeek V4 Pro 0813 | DeepSeek | 2026-08-12 | 1.0M | 2026-08-12 |
| 09 | SpaceXAI: Grok 4.6 | xAI | 2026-08-12 | 500K | 2026-08-12 |
| 10 | Meta: Muse Glimmer 30B | Meta | 2026-08-09 | 131K | 2026-08-09 |
| 11 | Meta: Muse Spark 1.2 | Meta | 2026-08-05 | 1.0M | 2026-08-05 |
| 12 | Qwen: Qwen3.8 Max | Alibaba | 2026-08-03 | 1.0M | 2026-08-03 |
| 13 | DeepSeek: DeepSeek V4 Flash 0731 | DeepSeek | 2026-07-31 | 1.0M | 2026-07-31 |
| 14 | Claude Opus 5 | Anthropic | 2026-07-24 | 1.0M | 2026-07-24 |
| 15 | Thinking Machines: Inkling | Thinking Machines | 2026-07-17 | 1.0M | 2026-07-17 |
| 16 | MoonshotAI: Kimi K3 | Moonshot | 2026-07-16 | 1.0M | 2026-07-16 |
| 17 | Meta: Muse Spark 1.1 | Meta | 2026-07-16 | 1.0M | 2026-07-16 |
| 18 | OpenAI: GPT-5.6 Terra Pro | OpenAI | 2026-07-09 | 1.1M | 2026-07-09 |
| 19 | OpenAI: GPT-5.6 Sol Pro | OpenAI | 2026-07-09 | 1.1M | 2026-07-09 |
| 20 | OpenAI: GPT-5.6 Sol | OpenAI | 2026-07-09 | 1.1M | 2026-07-09 |
| 21 | SpaceXAI: Grok 4.5 | xAI | 2026-07-08 | 500K | 2026-07-08 |
| 22 | GPT-5.6 Terra | OpenAI | 2026-07-08 | 400K | 2026-07-08 |
| 23 | Tencent: Hy3 | Tencent | 2026-07-06 | 262K | 2026-07-06 |
| 24 | Anthropic: Claude Sonnet 5 | Anthropic | 2026-06-30 | 1.0M | 2026-06-30 |
| 25 | Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | 2026-06-30 | 66K | 2026-06-30 | |
| 26 | Sakana: Fugu Ultra | Sakana | 2026-06-24 | 1.0M | 2026-06-24 |
| 27 | Google: Nano Banana 2 (Gemini 3.1 Flash Image) | 2026-06-18 | 131K | 2026-06-18 | |
| 28 | Google: Nano Banana Pro (Gemini 3 Pro Image) | 2026-06-18 | 66K | 2026-06-18 | |
| 29 | Z.ai: GLM 5.2 | Z.ai | 2026-06-16 | 1.0M | 2026-06-16 |
| 30 | Qwen 3 Max | Alibaba | 2026-06-15 | 262K | 2026-06-15 |
| 31 | Grok 4 | xAI | 2026-06-11 | 256K | 2026-06-11 |
| 32 | Anthropic: Claude Fable 5 | Anthropic | 2026-06-09 | 1.0M | 2026-06-09 |
| 33 | NVIDIA: Nemotron 3 Ultra | NVIDIA | 2026-06-04 | 262K | 2026-06-04 |
| 34 | MiniMax: MiniMax M3 | MiniMax | 2026-05-31 | 1.0M | 2026-05-31 |
| 35 | StepFun: Step 3.7 Flash | StepFun | 2026-05-28 | 262K | 2026-05-28 |
| 36 | Anthropic: Claude Opus 4.8 (Fast) | Anthropic | 2026-05-27 | 1.0M | 2026-05-27 |
| 37 | Anthropic: Claude Opus 4.8 | Anthropic | 2026-05-27 | 1.0M | 2026-05-27 |
| 38 | Qwen: Qwen3.7 Max | Alibaba | 2026-05-21 | 1.0M | 2026-05-21 |
| 39 | SpaceXAI: Grok Build 0.1 | xAI | 2026-05-20 | 256K | 2026-05-20 |
| 40 | Claude 4.5 Opus | Anthropic | 2026-05-19 | 500K | 2026-05-19 |
| 41 | Mistral Large 3 | Mistral | 2026-05-08 | 256K | 2026-05-08 |
| 42 | Google: Gemini 3.1 Flash Lite | 2026-05-07 | 1.0M | 2026-05-07 | |
| 43 | Mistral: Mistral Medium 3.5 | Mistral | 2026-04-30 | 262K | 2026-04-30 |
| 44 | Qwen: Qwen3.5 Plus 2026-04-20 | Alibaba | 2026-04-27 | 1.0M | 2026-04-27 |
| 45 | Qwen: Qwen3.6 Max Preview | Alibaba | 2026-04-27 | 262K | 2026-04-27 |
| 46 | OpenAI: GPT-5.5 Pro | OpenAI | 2026-04-24 | 1.1M | 2026-04-24 |
| 47 | OpenAI: GPT-5.5 | OpenAI | 2026-04-24 | 1.1M | 2026-04-24 |
| 48 | DeepSeek: DeepSeek V4 Pro 0423 | DeepSeek | 2026-04-24 | 1.0M | 2026-04-24 |
| 49 | DeepSeek: DeepSeek V4 Flash | DeepSeek | 2026-04-24 | 1.0M | 2026-04-24 |
| 50 | OpenAI: GPT-5.4 Image 2 | OpenAI | 2026-04-21 | 272K | 2026-04-21 |
| 51 | MoonshotAI: Kimi K2.6 | Moonshot | 2026-04-20 | 262K | 2026-04-20 |
| 52 | Qwen: Qwen3.6 Plus | Alibaba | 2026-04-02 | 1.0M | 2026-04-02 |
| 53 | Z.ai: GLM 5V Turbo | Z.ai | 2026-04-01 | 203K | 2026-04-01 |
| 54 | SpaceXAI: Grok 4.20 | xAI | 2026-03-31 | 2.0M | 2026-03-31 |
| 55 | OpenAI: GPT-5.4 | OpenAI | 2026-03-05 | 1.1M | 2026-03-05 |
| 56 | Qwen: Qwen3.5 397B A17B | Alibaba | 2026-02-16 | 262K | 2026-02-16 |
| 57 | Z.ai: GLM 5 | Z.ai | 2026-02-11 | 205K | 2026-02-11 |
| 58 | Qwen: Qwen3 Max Thinking | Alibaba | 2026-02-09 | 262K | 2026-02-09 |
| 59 | Mistral: Mistral Large 3 2512 | Mistral | 2025-12-01 | 262K | 2025-12-01 |
How to judge a model in its first weeks
New does not mean better, and the first two weeks after a launch are the least reliable time to evaluate any model. Launch-day benchmark numbers come from the lab, using prompting and scaffolding it chose, with no independent replication yet. Independent harnesses usually report lower figures a few weeks later, and that gap is normal rather than dishonest.
Serving conditions also shift after release. Providers tune quantisation, batching and routing under real load, so the model available on day one is not always identical in behaviour to the one available a month later. Independently-run indices measured against production endpoints are more trustworthy than launch-post tables for exactly this reason.
There is a real advantage to newness, though: training cutoffs. A recently trained model knows about libraries, APIs and events that an older one does not, which matters enormously for coding against fast-moving frameworks. If your work involves recent tooling, recency can outweigh a few benchmark points.
The failure mode to avoid is chasing every release. Migration has costs — prompts that were tuned for one model rarely transfer cleanly, tool-calling behaviour shifts, and output formatting changes in ways that break downstream parsing. A sensible cadence is to evaluate new frontier releases against your own test set, and only migrate when the improvement is clearly larger than the switching cost.
This page orders models by public release date so you can see how fresh the current leader actually is. Rumoured and unreleased models are tracked separately and marked as such — those entries reflect credible public reporting, not confirmed specifications.
More on how these numbers are produced in the methodology and what each evaluation measures in the benchmark guide. Spotted a score that disagrees with its source? Tell us.
Frequently asked
- What is the newest LLM released?
- The top of this table shows the most recently released frontier model tracked here. Release cadence has accelerated — expect a new flagship every 4–8 weeks from at least one lab.