Release timeline

The newest LLMs, freshest first

Ranked by release date. New flagship models drop every few weeks — this list is the fastest way to see what just shipped from OpenAI, Anthropic, Google, xAI, DeepSeek, Meta, Kimi and others.

Category leader

#01Alibaba

Qwen: Qwen3.8 Max (0902).

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...

Released
2026-09-03

Ranking

N=55
#ModelProviderReleasedCtxReleased
01Qwen: Qwen3.8 Max (0902)Alibaba2026-09-031.0M2026-09-03
02Meta: Muse Spark 1.3Meta2026-09-021.0M2026-09-02
03Google: Gemini 3.8 FlashGoogle2026-09-021.0M2026-09-02
04Anthropic: Claude Fable 5.1Anthropic2026-09-011.0M2026-09-01
05Z.ai: GLM 5.3Z.ai2026-08-181.3M2026-08-18
06Qwen: Qwen3.8 2.4T A95BAlibaba2026-08-121.0M2026-08-12
07DeepSeek: DeepSeek V4 Pro 0813DeepSeek2026-08-121.0M2026-08-12
08SpaceXAI: Grok 4.6xAI2026-08-12500K2026-08-12
09Meta: Muse Glimmer 30BMeta2026-08-09131K2026-08-09
10DeepSeek: DeepSeek V4 Flash 0731DeepSeek2026-07-311.0M2026-07-31
11Claude Opus 5Anthropic2026-07-241.0M2026-07-24
12MoonshotAI: Kimi K3Moonshot2026-07-161.0M2026-07-16
13Meta: Muse Spark 1.1Meta2026-07-161.0M2026-07-16
14OpenAI: GPT-5.6 Terra ProOpenAI2026-07-091.1M2026-07-09
15OpenAI: GPT-5.6 Sol ProOpenAI2026-07-091.1M2026-07-09
16OpenAI: GPT-5.6 SolOpenAI2026-07-091.1M2026-07-09
17SpaceXAI: Grok 4.5xAI2026-07-08500K2026-07-08
18GPT-5.6 TerraOpenAI2026-07-08400K2026-07-08
19Tencent: Hy3Tencent2026-07-06262K2026-07-06
20Anthropic: Claude Sonnet 5Anthropic2026-06-301.0M2026-06-30
21Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)Google2026-06-3066K2026-06-30
22Sakana: Fugu UltraSakana2026-06-241.0M2026-06-24
23Google: Nano Banana 2 (Gemini 3.1 Flash Image)Google2026-06-18131K2026-06-18
24Google: Nano Banana Pro (Gemini 3 Pro Image)Google2026-06-1866K2026-06-18
25Z.ai: GLM 5.2Z.ai2026-06-161.0M2026-06-16
26Qwen 3 MaxAlibaba2026-06-15262K2026-06-15
27Grok 4xAI2026-06-11256K2026-06-11
28Anthropic: Claude Fable 5Anthropic2026-06-091.0M2026-06-09
29MiniMax: MiniMax M3MiniMax2026-05-311.0M2026-05-31
30StepFun: Step 3.7 FlashStepFun2026-05-28262K2026-05-28
31Anthropic: Claude Opus 4.8 (Fast)Anthropic2026-05-271.0M2026-05-27
32Anthropic: Claude Opus 4.8Anthropic2026-05-271.0M2026-05-27
33Qwen: Qwen3.7 MaxAlibaba2026-05-211.0M2026-05-21
34SpaceXAI: Grok Build 0.1xAI2026-05-20256K2026-05-20
35Claude 4.5 OpusAnthropic2026-05-19500K2026-05-19
36Mistral Large 3Mistral2026-05-08256K2026-05-08
37Google: Gemini 3.1 Flash LiteGoogle2026-05-071.0M2026-05-07
38Mistral: Mistral Medium 3.5Mistral2026-04-30262K2026-04-30
39Llama 4 405BMeta2026-04-30128K2026-04-30
40Qwen: Qwen3.5 Plus 2026-04-20Alibaba2026-04-271.0M2026-04-27
41Qwen: Qwen3.6 Max PreviewAlibaba2026-04-27262K2026-04-27
42OpenAI: GPT-5.5 ProOpenAI2026-04-241.1M2026-04-24
43OpenAI: GPT-5.5OpenAI2026-04-241.1M2026-04-24
44DeepSeek: DeepSeek V4 Pro 0423DeepSeek2026-04-241.0M2026-04-24
45DeepSeek: DeepSeek V4 FlashDeepSeek2026-04-241.0M2026-04-24
46OpenAI: GPT-5.4 Image 2OpenAI2026-04-21272K2026-04-21
47MoonshotAI: Kimi K2.6Moonshot2026-04-20262K2026-04-20
48Qwen: Qwen3.6 PlusAlibaba2026-04-021.0M2026-04-02
49Z.ai: GLM 5V TurboZ.ai2026-04-01203K2026-04-01
50SpaceXAI: Grok 4.20xAI2026-03-312.0M2026-03-31
51OpenAI: GPT-5.4OpenAI2026-03-051.1M2026-03-05
52Qwen: Qwen3.5 397B A17BAlibaba2026-02-16262K2026-02-16
53Z.ai: GLM 5Z.ai2026-02-11205K2026-02-11
54Qwen: Qwen3 Max ThinkingAlibaba2026-02-09262K2026-02-09
55Mistral: Mistral Large 3 2512Mistral2025-12-01262K2025-12-01

How to judge a model in its first weeks

New does not mean better, and the first two weeks after a launch are the least reliable time to evaluate any model. Launch-day benchmark numbers come from the lab, using prompting and scaffolding it chose, with no independent replication yet. Independent harnesses usually report lower figures a few weeks later, and that gap is normal rather than dishonest.

Serving conditions also shift after release. Providers tune quantisation, batching and routing under real load, so the model available on day one is not always identical in behaviour to the one available a month later. Independently-run indices measured against production endpoints are more trustworthy than launch-post tables for exactly this reason.

There is a real advantage to newness, though: training cutoffs. A recently trained model knows about libraries, APIs and events that an older one does not, which matters enormously for coding against fast-moving frameworks. If your work involves recent tooling, recency can outweigh a few benchmark points.

The failure mode to avoid is chasing every release. Migration has costs — prompts that were tuned for one model rarely transfer cleanly, tool-calling behaviour shifts, and output formatting changes in ways that break downstream parsing. A sensible cadence is to evaluate new frontier releases against your own test set, and only migrate when the improvement is clearly larger than the switching cost.

This page orders models by public release date so you can see how fresh the current leader actually is. Rumoured and unreleased models are tracked separately and marked as such — those entries reflect credible public reporting, not confirmed specifications.

More on how these numbers are produced in the methodology and what each evaluation measures in the benchmark guide. Spotted a score that disagrees with its source? Tell us.

Frequently asked

What is the newest LLM released?
The top of this table shows the most recently released frontier model tracked here. Release cadence has accelerated — expect a new flagship every 4–8 weeks from at least one lab.