SOTA Model — the live state-of-the-art AI leaderboard. Current SOTA: Anthropic: Claude Fable 5.1 by Anthropic.
Current state of the art · v4.0 · streaming
“Claude Fable 5.1 excels on SWE-bench Verified, Aider Polyglot, and Terminal-Bench, with elite LiveBench and LMArena performance for long-horizon coding and agent workflows.”
SOTA Elo
+19 vs #2 in the field
SOTA GPQA
Graduate-level science QA
Max Context
Token window at the top
Avg Output $
Across tracked frontier tier
Pick a lens
SOTA depends on the job.
The general SOTA above answers "what's the best model." These leaderboards answer "best for what."
Model Telemetry
| 01 | Anthropic: Claude Fable 5.1Anthropic SOTA | 1638 | 71 | 91.1 | 93.7 | 86.1 | 77.4 | 43.8 | 89.6 | 1M | $50.00 | 3 days ago |
| 02 | Qwen: Qwen3.8 Max (0902)Alibaba | 1619 | 68 | 89.4 | 91.5 | 78.7 | 69.3 | 39.7 | 84.9 | 1M | $6.00 | 1 day ago |
| 03 | 1627 | 68 | 89.7 | 92.6 | 80.3 | 71.8 | 42.6 | 85.1 | 500K | $6.00 | 23 days ago | |
| 04 | MoonshotAI: Kimi K3Moonshot | 1608 | 67 | 88.5 | 90.8 | 79.4 | 70.2 | 37.2 | 86.8 | 1.0M | $15.00 | 1 month ago |
| 05 | Z.ai: GLM 5.3Z.ai | 1604 | 66 | 87.9 | 89.7 | 77.8 | 70.6 | 34.8 | 86.3 | 1.3M | $4.40 | 17 days ago |
| 06 | 1598 | 65 | 87.6 | 89.9 | 76.9 | 68.9 | 36.1 | 83.7 | 1.0M | $4.25 | 2 days ago | |
| 07 | DeepSeek: DeepSeek V4 Pro 0813DeepSeek | 1591 | 65 | 87.1 | 88.8 | 76.4 | 67.5 | 33.5 | 85.6 | 1.0M | $3.36 | 23 days ago |
| 08 | Claude Opus 5Anthropic | 1609 | 67 | 89.0 | 91.8 | 81.6 | 72.7 | 40.6 | 87.1 | 1M | $25.00 | 1 month ago |
| 09 | OpenAI: GPT-5.6 SolOpenAI | 1601 | 66 | 88.6 | 90.6 | 78.8 | 72.1 | 41.4 | 84.5 | 1.1M | $10.00 | 1 month ago |
| 10 | Qwen: Qwen3.8 2.4T A95BAlibaba | 1588 | 64 | 86.7 | 88.4 | 74.8 | 65.2 | 32.8 | 82.4 | 1.0M | $6.00 | 23 days ago |
| 11 | OpenAI: GPT-5.5 ProOpenAI | 1596 | 65 | 87.8 | 90.4 | 77.4 | 70.3 | 40.1 | 82.8 | 1.1M | $180.00 | 4 months ago |
| 12 | Anthropic: Claude Fable 5Anthropic | 1594 | 65 | 88.1 | 90.1 | 80.2 | 71.6 | 39.5 | 86.2 | 1M | $50.00 | 2 months ago |
| 13 | 1585 | 64 | 86.9 | 89.8 | 76.8 | 67.7 | 37.9 | 82.6 | 500K | $6.00 | 1 month ago | |
| 14 | Z.ai: GLM 5.2Z.ai | 1574 | 62 | 85.6 | 87.6 | 73.9 | 64.8 | 31.4 | 82.2 | 1.0M | $3.04 | 2 months ago |
| 15 | MoonshotAI: Kimi K2.6Moonshot | 1579 | 63 | 85.9 | 88.2 | 75.8 | 65.9 | 32.6 | 84.1 | 262K | $4.00 | 4 months ago |
| 16 | Anthropic: Claude Opus 4.8Anthropic | 1583 | 64 | 86.5 | 89.1 | 78.6 | 69.1 | 36.3 | 84.8 | 1M | $25.00 | 3 months ago |
| 17 | Qwen: Qwen3.7 MaxAlibaba | 1567 | 60 | 84.2 | 86.5 | 71.2 | 61.8 | 29.7 | 79.5 | 1M | $4.42 | 3 months ago |
| 18 | DeepSeek: DeepSeek V4 Pro 0423DeepSeek | 1564 | 61 | 84.5 | 86.2 | 72.6 | 62.7 | 29.8 | 80.6 | 1.0M | $1.86 | 4 months ago |
| 19 | OpenAI: GPT-5.5OpenAI | 1572 | 62 | 85.4 | 88.1 | 74.1 | 66.2 | 34.6 | 80.1 | 1.1M | $30.00 | 4 months ago |
| 20 | 1559 | 59 | 83.6 | 85.8 | 70.8 | 61.4 | 30.4 | 77.8 | 2M | $2.50 | 5 months ago | |
| 21 | Z.ai: GLM 5Z.ai | 1554 | 58 | 82.7 | 84.7 | 69.7 | 60.3 | 27.9 | 78.9 | 205K | $1.92 | 6 months ago |
| 22 | Qwen: Qwen3 Max ThinkingAlibaba | 1548 | 57 | 81.9 | 84.1 | 68.4 | 58.7 | 26.8 | 77.4 | 262K | $3.90 | 6 months ago |
| 23 | Anthropic: Claude Sonnet 5Anthropic | 1539 | 62 | 79.0 | 85.5 | 77.6 | 62.7 | 38.7 | 84.0 | 1M | $10.00 | 2 months ago |
| 24 | 1442 | 63 | 72.8 | 80.5 | 64.6 | 49.4 | 15.2 | 74.6 | 131K | $1.10 | 26 days ago | |
| 25 | Qwen: Qwen3.6 PlusAlibaba | — | — | — | — | — | — | — | — | 1M | $1.95 | 5 months ago |
| 26 | Shieldstral 1.0 3BMistral AI | — | — | — | — | — | — | — | — | — | — | — |
| 27 | MiniMax H3MiniMax | — | — | — | — | — | — | — | — | — | — | — |
| 28 | GPT-transcribeOpenAI | — | — | — | — | — | — | — | — | — | — | — |
| 29 | GPT-5.6 TerraOpenAI | 1462 | — | — | 80.2 | — | — | — | — | 400K | $4.80 | 1 month ago |
| 30 | Qwen Image 3.0 ProAlibaba / Qwen | — | — | — | — | — | — | — | — | — | — | — |
| 31 | Qwen3.8-MaxAlibaba / Qwen | — | — | — | — | — | — | — | — | — | — | — |
| 32 | Kimi K3Moonshot | — | — | — | — | — | — | — | — | — | — | — |
| 33 | 1547 | 62 | 76.8 | 85.1 | 74.7 | 62.6 | 12.4 | 80.8 | 1.0M | $4.25 | 1 month ago | |
| 34 | Mistral: Mistral Medium 3.5Mistral | 1481 | 56 | 73.8 | 81.2 | 64.3 | 52.8 | 18.2 | 65.4 | 262K | $7.50 | 4 months ago |
| 35 | Tencent: Hy3Tencent | 1542 | 60 | 73.8 | 80.7 | 62.1 | 47.2 | 20.6 | 70.2 | 262K | $0.53 | 2 months ago |
| 36 | Google WeatherNext 3Google | — | — | — | — | — | — | — | — | — | — | — |
| 37 | SIMA 2Google DeepMind | — | — | — | — | — | — | — | — | — | — | — |
| 38 | Anthropic: Claude Opus 4.8 (Fast)Anthropic | 1518 | 49 | 78.8 | 86.7 | 72.8 | 60.9 | 14.3 | 79.8 | 1M | $50.00 | 3 months ago |
| 39 | Google: Gemini 3.8 FlashGoogle | 1543 | 63 | 80.0 | 85.9 | 76.1 | 61.5 | 39.8 | 80.4 | 1.0M | $3.75 | 2 days ago |
| 40 | StepFun: Step 3.7 FlashStepFun | — | — | — | — | — | — | — | — | 262K | $1.15 | 3 months ago |
| 41 | Qwen: Qwen3.5 397B A17BAlibaba | 1463 | 53 | 72.4 | 81.2 | 64.7 | 50.8 | 11.5 | 75.4 | 262K | $3.60 | 6 months ago |
| 42 | 1412 | 53 | 71.4 | 76.8 | 54.7 | 38.1 | 13.1 | 71.5 | 1M | $1.80 | 4 months ago | |
| 43 | DeepSeek: DeepSeek V4 Flash 0731DeepSeek | 1495 | 60 | 75.0 | 80.0 | 70.0 | 57.0 | 17.0 | 81.0 | 1.0M | $0.18 | 1 month ago |
| 44 | Gemini 3.6 FlashGoogle | — | — | — | — | — | — | — | — | — | — | — |
| 45 | Claude Mythos 5Anthropic | — | — | — | — | — | — | — | — | — | — | — |
| 46 | Gemini Robotics 2Google DeepMind | — | — | — | — | — | — | — | — | — | — | — |
| 47 | Google TimesFM 3Google | — | — | — | — | — | — | — | — | — | — | — |
| 48 | Grok 4xAI | 1402 | — | — | 74.1 | — | — | — | — | 256K | $8.00 | 2 months ago |
| 49 | Claude FableAnthropic | — | — | — | — | — | — | — | — | — | — | — |
| 50 | OpenAI: GPT-5.4 Image 2OpenAI | — | — | — | — | — | — | — | — | 272K | $15.00 | 4 months ago |
| 51 | 1478 | 62 | 73.5 | 81.9 | 63.5 | 52.4 | 19.8 | 77.4 | 256K | $2.00 | 3 months ago | |
| 52 | — | — | — | — | — | — | — | — | 66K | $12.00 | 2 months ago | |
| 53 | — | — | — | — | — | — | — | — | — | — | — | |
| 54 | DeepSeek: DeepSeek V4 FlashDeepSeek | 1415 | 52 | 62.0 | 70.0 | 53.0 | 38.0 | 7.0 | 70.0 | 1.0M | $0.28 | 4 months ago |
| 55 | — | — | — | — | — | — | — | — | 66K | $1.50 | 2 months ago | |
| 56 | MiniMax: MiniMax M3MiniMax | 1535 | 59 | 71.3 | 80.8 | 65.1 | 53.2 | 12.1 | 71.9 | 1.0M | $1.20 | 3 months ago |
| 57 | Mistral Large 3Mistral | 1358 | — | — | 66.4 | — | — | — | — | 256K | $6.00 | 4 months ago |
| 58 | OpenAI: GPT-5.4OpenAI | 1570 | 71 | 80.7 | 88.8 | 73.7 | 61.0 | 22.7 | 77.8 | 1.1M | $15.00 | 6 months ago |
| 59 | DeepSeek V4 Flash 0731DeepSeek | — | — | — | — | — | — | — | — | — | — | — |
| 60 | Qwen 3 MaxAlibaba | 1371 | — | — | 68.9 | — | — | — | — | 262K | $1.50 | 2 months ago |
| 61 | 1616 | 66 | 83.0 | 89.1 | 76.4 | 63.9 | 35.8 | 83.8 | 1.1M | $12.00 | 1 month ago | |
| 62 | Ling-3.0-FlashLing | — | — | — | — | — | — | — | — | — | — | — |
| 63 | Grok 4.6xAI | — | — | — | — | — | — | — | — | — | — | — |
| 64 | GPT-5.6OpenAI | — | — | — | — | — | — | — | — | — | — | — |
| 65 | Claude Mythos 5.1Anthropic | — | — | — | — | — | — | — | — | — | — | — |
| 66 | Sakana: Fugu UltraSakana | 1578 | 67 | 85.8 | 85.9 | 74.6 | 63.7 | 38.6 | 80.8 | 1M | $30.00 | 2 months ago |
| 67 | 1568 | 65 | 83.8 | 83.9 | 71.8 | 60.3 | 35.6 | 79.2 | 262K | $1.50 | 9 months ago | |
| 68 | 1447 | 55 | 71.1 | 77.9 | 64.1 | 48.5 | 26.9 | 76.1 | 203K | $4.00 | 5 months ago | |
| 69 | — | — | — | — | — | — | — | — | — | — | — | |
| 70 | Qwen: Qwen3.6 Max PreviewAlibaba | 1485 | 72 | 71.0 | 77.0 | 60.0 | 50.0 | 19.0 | 74.0 | 262K | $6.24 | 4 months ago |
| 71 | — | — | — | — | — | — | — | — | 1.0M | $1.50 | 4 months ago | |
| 72 | — | — | — | — | — | — | — | — | 131K | $3.00 | 2 months ago | |
| 73 | Gemini Robotics ER 2Google DeepMind | — | — | — | — | — | — | — | — | — | — | — |
| 74 | Claude 4.5 OpusAnthropic | 1458 | — | — | 81.9 | — | — | — | — | 500K | $24.00 | 3 months ago |
| 75 | OpenAI: GPT-5.6 Sol ProOpenAI | 1542 | 69 | 86.3 | 91.0 | 80.3 | 68.5 | 22.8 | 80.1 | 1.1M | $10.00 | 1 month ago |
| 76 | Llama 4 405BMeta | 1389 | — | — | 71.2 | — | — | — | — | 128K | $2.70 | 4 months ago |
| 77 | Qwen Drive 1.0Alibaba Qwen | — | — | — | — | — | — | — | — | — | — | — |
| 78 | Qwen 3.8Alibaba | — | — | — | — | — | — | — | — | — | — | — |
| 79 | Claude Gen-5Anthropic | — | — | — | — | — | — | — | — | — | — | — |
| 80 | Claude MythosAnthropic | — | — | — | — | — | — | — | — | — | — | — |
Anthropic · 3 days ago
Elo
1638
AA
71
GPQA
93.7
SWE
86
Term
77
ARC
44
LiveB
91
Ctx
1M
Alibaba · 1 day ago
Elo
1619
AA
68
GPQA
91.5
SWE
79
Term
69
ARC
40
LiveB
89
Ctx
1M
xAI · 23 days ago
Elo
1627
AA
68
GPQA
92.6
SWE
80
Term
72
ARC
43
LiveB
90
Ctx
500K
Moonshot · 1 month ago
Elo
1608
AA
67
GPQA
90.8
SWE
79
Term
70
ARC
37
LiveB
89
Ctx
1.0M
Z.ai · 17 days ago
Elo
1604
AA
66
GPQA
89.7
SWE
78
Term
71
ARC
35
LiveB
88
Ctx
1.3M
Meta · 2 days ago
Elo
1598
AA
65
GPQA
89.9
SWE
77
Term
69
ARC
36
LiveB
88
Ctx
1.0M
DeepSeek · 23 days ago
Elo
1591
AA
65
GPQA
88.8
SWE
76
Term
68
ARC
34
LiveB
87
Ctx
1.0M
Anthropic · 1 month ago
Elo
1609
AA
67
GPQA
91.8
SWE
82
Term
73
ARC
41
LiveB
89
Ctx
1M
OpenAI · 1 month ago
Elo
1601
AA
66
GPQA
90.6
SWE
79
Term
72
ARC
41
LiveB
89
Ctx
1.1M
Alibaba · 23 days ago
Elo
1588
AA
64
GPQA
88.4
SWE
75
Term
65
ARC
33
LiveB
87
Ctx
1.0M
OpenAI · 4 months ago
Elo
1596
AA
65
GPQA
90.4
SWE
77
Term
70
ARC
40
LiveB
88
Ctx
1.1M
Anthropic · 2 months ago
Elo
1594
AA
65
GPQA
90.1
SWE
80
Term
72
ARC
40
LiveB
88
Ctx
1M
xAI · 1 month ago
Elo
1585
AA
64
GPQA
89.8
SWE
77
Term
68
ARC
38
LiveB
87
Ctx
500K
Z.ai · 2 months ago
Elo
1574
AA
62
GPQA
87.6
SWE
74
Term
65
ARC
31
LiveB
86
Ctx
1.0M
Moonshot · 4 months ago
Elo
1579
AA
63
GPQA
88.2
SWE
76
Term
66
ARC
33
LiveB
86
Ctx
262K
Anthropic · 3 months ago
Elo
1583
AA
64
GPQA
89.1
SWE
79
Term
69
ARC
36
LiveB
87
Ctx
1M
Alibaba · 3 months ago
Elo
1567
AA
60
GPQA
86.5
SWE
71
Term
62
ARC
30
LiveB
84
Ctx
1M
DeepSeek · 4 months ago
Elo
1564
AA
61
GPQA
86.2
SWE
73
Term
63
ARC
30
LiveB
85
Ctx
1.0M
OpenAI · 4 months ago
Elo
1572
AA
62
GPQA
88.1
SWE
74
Term
66
ARC
35
LiveB
85
Ctx
1.1M
xAI · 5 months ago
Elo
1559
AA
59
GPQA
85.8
SWE
71
Term
61
ARC
30
LiveB
84
Ctx
2M
Z.ai · 6 months ago
Elo
1554
AA
58
GPQA
84.7
SWE
70
Term
60
ARC
28
LiveB
83
Ctx
205K
Alibaba · 6 months ago
Elo
1548
AA
57
GPQA
84.1
SWE
68
Term
59
ARC
27
LiveB
82
Ctx
262K
Anthropic · 2 months ago
Elo
1539
AA
62
GPQA
85.5
SWE
78
Term
63
ARC
39
LiveB
79
Ctx
1M
Meta · 26 days ago
Elo
1442
AA
63
GPQA
80.5
SWE
65
Term
49
ARC
15
LiveB
73
Ctx
131K
Alibaba · 5 months ago
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
1M
Mistral AI · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
MiniMax · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
OpenAI · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
OpenAI · 1 month ago
Elo
1462
AA
—
GPQA
80.2
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
400K
Alibaba / Qwen · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Alibaba / Qwen · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Moonshot · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Meta · 1 month ago
Elo
1547
AA
62
GPQA
85.1
SWE
75
Term
63
ARC
12
LiveB
77
Ctx
1.0M
Mistral · 4 months ago
Elo
1481
AA
56
GPQA
81.2
SWE
64
Term
53
ARC
18
LiveB
74
Ctx
262K
Tencent · 2 months ago
Elo
1542
AA
60
GPQA
80.7
SWE
62
Term
47
ARC
21
LiveB
74
Ctx
262K
Google · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Google DeepMind · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Anthropic · 3 months ago
Elo
1518
AA
49
GPQA
86.7
SWE
73
Term
61
ARC
14
LiveB
79
Ctx
1M
Google · 2 days ago
Elo
1543
AA
63
GPQA
85.9
SWE
76
Term
62
ARC
40
LiveB
80
Ctx
1.0M
StepFun · 3 months ago
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
262K
Alibaba · 6 months ago
Elo
1463
AA
53
GPQA
81.2
SWE
65
Term
51
ARC
12
LiveB
72
Ctx
262K
Alibaba · 4 months ago
Elo
1412
AA
53
GPQA
76.8
SWE
55
Term
38
ARC
13
LiveB
71
Ctx
1M
DeepSeek · 1 month ago
Elo
1495
AA
60
GPQA
80.0
SWE
70
Term
57
ARC
17
LiveB
75
Ctx
1.0M
Google · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Anthropic · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Google DeepMind · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Google · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
xAI · 2 months ago
Elo
1402
AA
—
GPQA
74.1
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
256K
Anthropic · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
OpenAI · 4 months ago
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
272K
xAI · 3 months ago
Elo
1478
AA
62
GPQA
81.9
SWE
64
Term
52
ARC
20
LiveB
74
Ctx
256K
Google · 2 months ago
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
66K
xAI · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
DeepSeek · 4 months ago
Elo
1415
AA
52
GPQA
70.0
SWE
53
Term
38
ARC
7
LiveB
62
Ctx
1.0M
Google · 2 months ago
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
66K
MiniMax · 3 months ago
Elo
1535
AA
59
GPQA
80.8
SWE
65
Term
53
ARC
12
LiveB
71
Ctx
1.0M
Mistral · 4 months ago
Elo
1358
AA
—
GPQA
66.4
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
256K
OpenAI · 6 months ago
Elo
1570
AA
71
GPQA
88.8
SWE
74
Term
61
ARC
23
LiveB
81
Ctx
1.1M
DeepSeek · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Alibaba · 2 months ago
Elo
1371
AA
—
GPQA
68.9
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
262K
OpenAI · 1 month ago
Elo
1616
AA
66
GPQA
89.1
SWE
76
Term
64
ARC
36
LiveB
83
Ctx
1.1M
Ling · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
xAI · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
OpenAI · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Anthropic · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Sakana · 2 months ago
Elo
1578
AA
67
GPQA
85.9
SWE
75
Term
64
ARC
39
LiveB
86
Ctx
1M
Mistral · 9 months ago
Elo
1568
AA
65
GPQA
83.9
SWE
72
Term
60
ARC
36
LiveB
84
Ctx
262K
Z.ai · 5 months ago
Elo
1447
AA
55
GPQA
77.9
SWE
64
Term
49
ARC
27
LiveB
71
Ctx
203K
NVIDIA · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Alibaba · 4 months ago
Elo
1485
AA
72
GPQA
77.0
SWE
60
Term
50
ARC
19
LiveB
71
Ctx
262K
Google · 4 months ago
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
1.0M
Google · 2 months ago
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
131K
Google DeepMind · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Anthropic · 3 months ago
Elo
1458
AA
—
GPQA
81.9
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
500K
OpenAI · 1 month ago
Elo
1542
AA
69
GPQA
91.0
SWE
80
Term
69
ARC
23
LiveB
86
Ctx
1.1M
Meta · 4 months ago
Elo
1389
AA
—
GPQA
71.2
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
128K
Alibaba Qwen · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Alibaba · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Anthropic · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Anthropic · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Fresh drops
Latest models.
Newest first — including rumored and announced models that aren't yet on the leaderboard.
- OpenAI: GPT-6 Astra· OpenAI
OpenAI · 2026-09-01 — Path to Astra: critical capabilities and frontier safeguards
Announcedannounced · 13h ago - Qwen: Qwen3.8 Max (0902)· AlibabaReleasedreleased · 1d ago
- Meta: Muse Spark 1.3· MetaReleasedreleased · 2d ago
- Google: Gemini 3.8 Flash· Google
r/singularity · 2026-09-02 · Introducing Gemini 3.8 Flash
Releasedreleased · 2d ago - Anthropic: Claude Fable 5.1· AnthropicReleasedreleased · 3d ago
- Tencent: Hy4 preview· TencentReleasedreleased · 8d ago
- Releasedreleased · 14d ago
- DeepSeek: DeepSeek V4 Flash Vision Exp· DeepSeek
OpenRouter · preview endpoint live
Previewpreview · 14d ago - Z.ai: GLM 5.3· Z.ai
r/LocalLLaMA · official Z.ai announcement on 2026-08-14
Releasedreleased · 17d ago - Qwen: Qwen3.8 27B· AlibabaReleasedreleased · 21d ago
- Google: Gemini 3.7 Flash· Google
Ars Technica · Google announcement on 2026-08-13
Releasedreleased · 22d ago - Qwen: Qwen3.8 2.4T A95B· AlibabaReleasedreleased · 23d ago
- DeepSeek: DeepSeek V4 Pro 0813· DeepSeekReleasedreleased · 23d ago
- SpaceXAI: Grok 4.6· xAIReleasedreleased · 23d ago
- Meta: Muse Glimmer 30B· Meta
TechCrunch · Meta open-weight Muse Glimmer launch
Releasedreleased · 26d ago
Score Sources
Aggregated by GPT-5.6 TerraEvery score in the table above is aggregated from these public leaderboards. Click any column header to sort the leaderboard; the ↗ icon opens the underlying source. Model names link to their OpenRouter page.
- LMArena
Chatbot Arena Elo · human head-to-head votes
- Artificial Analysis
Intelligence Index · aggregate of 10+ evals
- LiveBench
Contamination-free monthly benchmark
- GPQA Diamond
Grad-level science reasoning (198 Qs)
- SWE-bench Verified
Real GitHub issue fixes · human-audited
- Terminal-Bench
Agentic terminal tasks · Stanford / LAION
- ARC-AGI-2
Abstract reasoning · François Chollet
- Aider Polyglot
Multi-language code editing benchmark
- MMLU-Pro
Expanded multitask academic knowledge
- OpenRouter
Live model catalog · pricing · throughput
How to read this board
EditorialThe table above is a composite, not a poll. Each model's position comes from reconciling independent public evaluations — human preference votes on LMArena, the independently-run Artificial Analysis index, contamination-resistant LiveBench, graduate-level GPQA Diamond, and execution-graded coding benchmarks like SWE-bench Verified and Terminal-Bench. A model earns the top slot only when it leads across the widest set of those signals at once; winning a single benchmark is never enough.
Treat gaps of a few points as ties. Every evaluation on this page has a noise band of several points, and lab-reported figures regularly differ from independently-run ones — when they disagree, we weight the independent run. Below the top handful of models, scores are best read as estimates with the source attached, not as precise measurements.
"SOTA" is also time-stamped. The current leader, Anthropic: Claude Fable 5.1, has held the position for roughly 4 days since its public release. The frontier typically turns over every few weeks, so the date under the hero matters as much as the name: a SOTA claim from last month is a historical fact, not a current one.
If you are choosing a model rather than watching the race, don't start here — start with the job. The use-case leaderboards for coding, reasoning and agentic work re-weight the same underlying scores for each workload, and the methodology page documents every weight and guardrail.
FAQ
Common questions.
- What is the SOTA LLM right now?
- The current state-of-the-art LLM on our aggregated leaderboard is shown in the hero at the top of this page. Rankings are refreshed every couple of hours from LMArena, Artificial Analysis, LiveBench, GPQA, SWE-bench, Terminal-Bench and more.
- What does SOTA mean in AI?
- SOTA stands for state of the art — the model that currently leads on the benchmarks a task cares about. For a general SOTA we aggregate across every major public evaluation; for task-specific SOTA see /coding, /reasoning, or /agentic.
- Which LLM is best for coding?
- The dedicated coding leaderboard at /coding ranks models by a weighted composite of SWE-bench Verified, Aider Polyglot, Terminal-Bench and LiveBench — the benchmarks that actually predict developer productivity.
- How often is this leaderboard updated?
- Every ~2 hours. The exact last-refresh timestamp is shown under the H1 and in the footer of every page.
Signal Stream
24 items- r/LocalLLaMA· 2h ago
Qwen3.8 27b for agentic coding and next .... what?
First, I'd like to thank the Qwen and Unsloth teams for the Qwen3.8 27b UD Q4_K_XL. Fits the poor 24GB of 3090 VRAM with 100k context at Q8 and works phenomenally well! Imho if theres anything that can threaten Anthropic/OpenAI profits is not another frontier model but actually these small ones you can run fast locally that can do 80..90% of mundane work for hours without paying a single dolla
- r/LocalLLaMA· 2h ago
I've found myself using Local LLM's like 3D printers.
Anyone who has a 3D printer and get use of it finds it incredibly useful for those odd jobs around the house, a missing bracket, a cable router, steam deck holder and so on. In the past if I was missing an app or useful software, a game I'd do the lazy thing, even though I can and have coded in the past, its "effort" I'll just go and buy or download the latest and greatest. Earli
- r/LocalLLaMA· 2h ago
The OpenAI Huggingface incident from an agents POV
Full credits to @artificialisabel from X!   submitted by   /u/iPingWine [link]   [comments]
- Hacker News· 3h ago
Could Anthropic have solved Navier–Stokes?
Article URL: https://twitter.com/ElliotGlazer/status/2096076054133952516 Comments URL: https://news.ycombinator.com/item?id=49573480 Points: 1 # Comments: 0
- Hacker News· 3h ago
Claude Fable 5.1 and Mythos 5.1: The System Card
Article URL: https://thezvi.substack.com/p/claude-fable-51-and-mythos-51-the Comments URL: https://news.ycombinator.com/item?id=49573407 Points: 1 # Comments: 0
- Hacker News· 4h ago
Ask HN: Replicate ChatGPT/Anthropic Voice Mode?
Hi, I'm interested in how to have my own phone number where I can call in and get a similar experience to ChatGPT voice mode or Anthropic voice mode. Anyone have experience setting this up? I've got something rudimentary setup with Parakeet for STT, Elevenlabs for TTS, Claude for the brains, and Twilio for the phone number handling. It would have impressed someone in 2024 but... Comments URL: http
- r/LocalLLaMA· 5h ago
How to tune llama.cpp codebase to add custom supported commands for AMD Vulkan?
Looking for any guidance regarding this, I want to tune my existing llama.cpp codebase so I can potentially implement custom supported commands (not random commands) for my existing amd vulkan setup, I don’t have plans to switch to Linux, I’m using window 11, so ROCm isn’t supported here The reason I’m doing this is because I’m looking forward to optimise the existing configuration, so I can poten
- r/LocalLLaMA· 5h ago
which model is good for detecting deflection?
I want the answers generated by frontier LLMs or base model LLM answers to be reviewed by some uncensored or abliterated small model. The job is this model (preferably small model) is just to detect deflection in the answers. The problem I am facing is uncensored SLM usually agrees on everything we give input. So the generated answer is also input for it and system prompt is input too.   submi
- Hacker News· 5h ago
Aegis – Inline security sidecar and eBPF sandbox for LLM agents
Article URL: https://aegiscruc.io Comments URL: https://news.ycombinator.com/item?id=49573010 Points: 1 # Comments: 0
- r/LocalLLaMA· 5h ago
Qwen3.8-27B beat the Wikipedia game in 6 clicks.
Used qwen3.8-27b in Opencode to make this silly mini-game because I'm not sober: ``` We are going to play a game, it will be the Wikipedia game. The Wikipedia game has the following rules: You will have a Wikipedia article set as a starting point. You will have a Wikipedia article set as an ending point. Your objective is to reach the the end point, which is an article completely separate from
- r/LocalLLaMA· 5h ago
AMD unveils Threadripper Halo Station
AMD Threadripper Halo Station CPU Ryzen Threadripper PRO 9995WX (Zen 5, "Shimada Peak") 96 cores / 192 threads Up to 5.4 GHz boost 384 MB L3 cache 350 W TDP 8-channel DDR5 128 PCIe 5.0 lanes System Memory 2 TB DDR5 (as shown at IFA) Accelerators 2 x Liquid Cooled AMD Instinct MI350P (CDNA 4) 144 GB HBM3E per card, up to 4 TB/s per card 288 GB total HBM3E Up to 600 W TBP per card PCIe 5.0
- Hacker News· 5h ago
GPT-6 Astra in code review: Gains, privacy, and cost
Article URL: https://www.coderabbit.ai/blog/gpt-6-astra-code-review-evaluation Comments URL: https://news.ycombinator.com/item?id=49572875 Points: 6 # Comments: 1
- r/singularity· 5h ago
Why would Google or NAVER partner with an independent AI web index?
Companies such as Parallel and Keenable are building new web indexes designed specifically for AI agents. Google has already integrated Parallel despite owning one of the world’s largest search indexes. Why would Google or NAVER partner with these companies instead of building the technology internally? Does this validate independent AI search as a major new category, or will these startups eventu
- r/singularity· 5h ago
One Shot Galaxy
Same prompt on Opus 4.8, Fable 5.1, and GPT 6 Astra.   submitted by   /u/LightningMcLovin [link]   [comments]
- r/LocalLLaMA· 5h ago
Chalk one up for the frontier model...
I just spent the last 2 hours of my life on a Friday night debugging a strange error in a prod CLI app. EF core was receive a readonlyspan during a Contains query. Normally, this query converted to a WHERE [col] IN (...) , but for some reason, after an update, it started choking, despite no code change. The same exact code runs in a separate website docker image fine, no problem. I put Qwen 3.8 27
- r/singularity· 6h ago
Fat Little Boy in an LLM Candy-shop
That’s me. Lately, I’ve been a fat little boy in the LLM candy-shop. After absolutely gorging myself on the Chinese sweets, K3, GLM 5.3, and Qwen 3.8 max, Elon lured me back to America with Grok 4.6 Panda Express. Then just the other day, boom. Zuck comes out of left field with Cherry Coke aka Muse Spark 1.3. Today, Asstra drops and in my gluttony, I’m using it to orchestrate Muse Spark workers. A
- r/singularity· 6h ago
AA Intelligence Index Changes
"Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming"   submitted by   /u/poigre [link]   [comments]
- r/singularity· 6h ago
Anthropic Possibly Tackles Its First Millennium Prize Problem
  submitted by   /u/ResultBackground2450 [link]   [comments]
- r/LocalLLaMA· 7h ago
RTX 4090 48GB longevity
Modified 4090 48GB has been out for a while. I remember a lot of people were buying them at the time. A lot of people were also complaining that they are meant to fail, that they scam etc. I have a few questions to people people who bought these. How is longevity of these cards? Do they still work without issues? Any failure rate? Do they use the same Nvidia drivers that regular 4090 or 4090D uses
- Hacker News· 8h ago
Claude Fable 5.1 vs. GPT-6 Astra, who wins on the 3D modeling?
Article URL: https://github.com/PhiloLabs/fable51-worlds/tree/main Comments URL: https://news.ycombinator.com/item?id=49572129 Points: 2 # Comments: 2
- r/singularity· 8h ago
AI will turn every profession into an art form instead of something to be productive on
I feel like a lot of the AI usage is opportunistic, we use it just because we can. Me first, that was my reasoning for using AI in a lot of work. But let's imagine we hit a cap of "infinite productivity", then will it really matter if we produce hand-made or using AI ? All needs everywhere will be fulfilled, for free, in every way possible. We won't feel more competitive by using
- r/LocalLLaMA· 8h ago
Instructions working well for qwen3.8
Important context: this is about preserve_thinking false stacks and makes no sense if you don't have that working end-to-end with your harness and llamacpp backend already sending reasoning_content and removing it. This is a bit tricky config-wise in llamacpp and your harness and not the default. I'm assuming the reader here already has a lot of prior knowledge. My entire goal is always to
- Hacker News· 8h ago
GPT-6 Astra: First Experience (high)
Article URL: https://chatgpt.com/share/6a9b671d-b780-83ea-bd5d-b0ea39f7d9b3 Comments URL: https://news.ycombinator.com/item?id=49571948 Points: 2 # Comments: 1
- r/singularity· 8h ago
GPT 6 Astra debuts with a 350 point lead on VoxelBench
  submitted by   /u/LightVelox [link]   [comments]