SOTA Model — the live state-of-the-art AI leaderboard. Current SOTA: Anthropic: Claude Fable 5.1 by Anthropic.

Current state of the art · v4.0 · streaming

SOTAAnthropic· 3 days ago
Anthropic: Claude Fable 5.1.
Field note#01 / SOTA

Claude Fable 5.1 excels on SWE-bench Verified, Aider Polyglot, and Terminal-Bench, with elite LiveBench and LMArena performance for long-horizon coding and agent workflows.

Provider
Anthropic
Released
3 days ago
01Anthropic: Claude Fable 5.11638
02Qwen: Qwen3.8 Max (0902)1619
03SpaceXAI: Grok 4.61627
04MoonshotAI: Kimi K31608
05Z.ai: GLM 5.31604
06Meta: Muse Spark 1.31598
07DeepSeek: DeepSeek V4 Pro 08131591
08Claude Opus 51609
09OpenAI: GPT-5.6 Sol1601
10Qwen: Qwen3.8 2.4T A95B1588

SOTA Elo

1,638

+19 vs #2 in the field

SOTA GPQA

93.7%

Graduate-level science QA

Max Context

1M

Token window at the top

Avg Output $

12.56/Mtok

Across tracked frontier tier

Pick a lens

SOTA depends on the job.

The general SOTA above answers "what's the best model." These leaderboards answer "best for what."

Model Telemetry

N=80
01
Anthropic: Claude Fable 5.1SOTA

Anthropic · 3 days ago

Elo

1638

AA

71

GPQA

93.7

SWE

86

Term

77

ARC

44

LiveB

91

Ctx

1M

02
Qwen: Qwen3.8 Max (0902)

Alibaba · 1 day ago

Elo

1619

AA

68

GPQA

91.5

SWE

79

Term

69

ARC

40

LiveB

89

Ctx

1M

03
SpaceXAI: Grok 4.6

xAI · 23 days ago

Elo

1627

AA

68

GPQA

92.6

SWE

80

Term

72

ARC

43

LiveB

90

Ctx

500K

04
MoonshotAI: Kimi K3

Moonshot · 1 month ago

Elo

1608

AA

67

GPQA

90.8

SWE

79

Term

70

ARC

37

LiveB

89

Ctx

1.0M

05
Z.ai: GLM 5.3

Z.ai · 17 days ago

Elo

1604

AA

66

GPQA

89.7

SWE

78

Term

71

ARC

35

LiveB

88

Ctx

1.3M

06
Meta: Muse Spark 1.3

Meta · 2 days ago

Elo

1598

AA

65

GPQA

89.9

SWE

77

Term

69

ARC

36

LiveB

88

Ctx

1.0M

07
DeepSeek: DeepSeek V4 Pro 0813

DeepSeek · 23 days ago

Elo

1591

AA

65

GPQA

88.8

SWE

76

Term

68

ARC

34

LiveB

87

Ctx

1.0M

08
Claude Opus 5

Anthropic · 1 month ago

Elo

1609

AA

67

GPQA

91.8

SWE

82

Term

73

ARC

41

LiveB

89

Ctx

1M

09
OpenAI: GPT-5.6 Sol

OpenAI · 1 month ago

Elo

1601

AA

66

GPQA

90.6

SWE

79

Term

72

ARC

41

LiveB

89

Ctx

1.1M

10
Qwen: Qwen3.8 2.4T A95B

Alibaba · 23 days ago

Elo

1588

AA

64

GPQA

88.4

SWE

75

Term

65

ARC

33

LiveB

87

Ctx

1.0M

11
OpenAI: GPT-5.5 Pro

OpenAI · 4 months ago

Elo

1596

AA

65

GPQA

90.4

SWE

77

Term

70

ARC

40

LiveB

88

Ctx

1.1M

12
Anthropic: Claude Fable 5

Anthropic · 2 months ago

Elo

1594

AA

65

GPQA

90.1

SWE

80

Term

72

ARC

40

LiveB

88

Ctx

1M

13
SpaceXAI: Grok 4.5

xAI · 1 month ago

Elo

1585

AA

64

GPQA

89.8

SWE

77

Term

68

ARC

38

LiveB

87

Ctx

500K

14
Z.ai: GLM 5.2

Z.ai · 2 months ago

Elo

1574

AA

62

GPQA

87.6

SWE

74

Term

65

ARC

31

LiveB

86

Ctx

1.0M

15
MoonshotAI: Kimi K2.6

Moonshot · 4 months ago

Elo

1579

AA

63

GPQA

88.2

SWE

76

Term

66

ARC

33

LiveB

86

Ctx

262K

16
Anthropic: Claude Opus 4.8

Anthropic · 3 months ago

Elo

1583

AA

64

GPQA

89.1

SWE

79

Term

69

ARC

36

LiveB

87

Ctx

1M

17
Qwen: Qwen3.7 Max

Alibaba · 3 months ago

Elo

1567

AA

60

GPQA

86.5

SWE

71

Term

62

ARC

30

LiveB

84

Ctx

1M

18
DeepSeek: DeepSeek V4 Pro 0423

DeepSeek · 4 months ago

Elo

1564

AA

61

GPQA

86.2

SWE

73

Term

63

ARC

30

LiveB

85

Ctx

1.0M

19
OpenAI: GPT-5.5

OpenAI · 4 months ago

Elo

1572

AA

62

GPQA

88.1

SWE

74

Term

66

ARC

35

LiveB

85

Ctx

1.1M

20
SpaceXAI: Grok 4.20

xAI · 5 months ago

Elo

1559

AA

59

GPQA

85.8

SWE

71

Term

61

ARC

30

LiveB

84

Ctx

2M

21
Z.ai: GLM 5

Z.ai · 6 months ago

Elo

1554

AA

58

GPQA

84.7

SWE

70

Term

60

ARC

28

LiveB

83

Ctx

205K

22
Qwen: Qwen3 Max Thinking

Alibaba · 6 months ago

Elo

1548

AA

57

GPQA

84.1

SWE

68

Term

59

ARC

27

LiveB

82

Ctx

262K

23
Anthropic: Claude Sonnet 5

Anthropic · 2 months ago

Elo

1539

AA

62

GPQA

85.5

SWE

78

Term

63

ARC

39

LiveB

79

Ctx

1M

24
Meta: Muse Glimmer 30B

Meta · 26 days ago

Elo

1442

AA

63

GPQA

80.5

SWE

65

Term

49

ARC

15

LiveB

73

Ctx

131K

25
Qwen: Qwen3.6 Plus

Alibaba · 5 months ago

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

1M

26
Shieldstral 1.0 3B

Mistral AI ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

27
MiniMax H3

MiniMax ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

28
GPT-transcribe

OpenAI ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

29
GPT-5.6 Terra

OpenAI · 1 month ago

Elo

1462

AA

GPQA

80.2

SWE

Term

ARC

LiveB

Ctx

400K

30
Qwen Image 3.0 Pro

Alibaba / Qwen ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

31
Qwen3.8-Max

Alibaba / Qwen ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

32
Kimi K3

Moonshot ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

33
Meta: Muse Spark 1.1

Meta · 1 month ago

Elo

1547

AA

62

GPQA

85.1

SWE

75

Term

63

ARC

12

LiveB

77

Ctx

1.0M

34
Mistral: Mistral Medium 3.5

Mistral · 4 months ago

Elo

1481

AA

56

GPQA

81.2

SWE

64

Term

53

ARC

18

LiveB

74

Ctx

262K

35
Tencent: Hy3

Tencent · 2 months ago

Elo

1542

AA

60

GPQA

80.7

SWE

62

Term

47

ARC

21

LiveB

74

Ctx

262K

36
Google WeatherNext 3

Google ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

37
SIMA 2

Google DeepMind ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

38
Anthropic: Claude Opus 4.8 (Fast)

Anthropic · 3 months ago

Elo

1518

AA

49

GPQA

86.7

SWE

73

Term

61

ARC

14

LiveB

79

Ctx

1M

39
Google: Gemini 3.8 Flash

Google · 2 days ago

Elo

1543

AA

63

GPQA

85.9

SWE

76

Term

62

ARC

40

LiveB

80

Ctx

1.0M

40
StepFun: Step 3.7 Flash

StepFun · 3 months ago

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

262K

41
Qwen: Qwen3.5 397B A17B

Alibaba · 6 months ago

Elo

1463

AA

53

GPQA

81.2

SWE

65

Term

51

ARC

12

LiveB

72

Ctx

262K

42
Qwen: Qwen3.5 Plus 2026-04-20

Alibaba · 4 months ago

Elo

1412

AA

53

GPQA

76.8

SWE

55

Term

38

ARC

13

LiveB

71

Ctx

1M

43
DeepSeek: DeepSeek V4 Flash 0731

DeepSeek · 1 month ago

Elo

1495

AA

60

GPQA

80.0

SWE

70

Term

57

ARC

17

LiveB

75

Ctx

1.0M

44
Gemini 3.6 Flash

Google ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

45
Claude Mythos 5

Anthropic ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

46
Gemini Robotics 2

Google DeepMind ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

47
Google TimesFM 3

Google ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

48
Grok 4

xAI · 2 months ago

Elo

1402

AA

GPQA

74.1

SWE

Term

ARC

LiveB

Ctx

256K

49
Claude Fable

Anthropic ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

50
OpenAI: GPT-5.4 Image 2

OpenAI · 4 months ago

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

272K

51
SpaceXAI: Grok Build 0.1

xAI · 3 months ago

Elo

1478

AA

62

GPQA

81.9

SWE

64

Term

52

ARC

20

LiveB

74

Ctx

256K

52
Google: Nano Banana Pro (Gemini 3 Pro Image)

Google · 2 months ago

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

66K

53
Grok Imagine 1.5

xAI ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

54
DeepSeek: DeepSeek V4 Flash

DeepSeek · 4 months ago

Elo

1415

AA

52

GPQA

70.0

SWE

53

Term

38

ARC

7

LiveB

62

Ctx

1.0M

55
Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)

Google · 2 months ago

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

66K

56
MiniMax: MiniMax M3

MiniMax · 3 months ago

Elo

1535

AA

59

GPQA

80.8

SWE

65

Term

53

ARC

12

LiveB

71

Ctx

1.0M

57
Mistral Large 3

Mistral · 4 months ago

Elo

1358

AA

GPQA

66.4

SWE

Term

ARC

LiveB

Ctx

256K

58
OpenAI: GPT-5.4

OpenAI · 6 months ago

Elo

1570

AA

71

GPQA

88.8

SWE

74

Term

61

ARC

23

LiveB

81

Ctx

1.1M

59
DeepSeek V4 Flash 0731

DeepSeek ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

60
Qwen 3 Max

Alibaba · 2 months ago

Elo

1371

AA

GPQA

68.9

SWE

Term

ARC

LiveB

Ctx

262K

61
OpenAI: GPT-5.6 Terra Pro

OpenAI · 1 month ago

Elo

1616

AA

66

GPQA

89.1

SWE

76

Term

64

ARC

36

LiveB

83

Ctx

1.1M

62
Ling-3.0-Flash

Ling ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

63
Grok 4.6

xAI ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

64
GPT-5.6

OpenAI ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

65
Claude Mythos 5.1

Anthropic ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

66
Sakana: Fugu Ultra

Sakana · 2 months ago

Elo

1578

AA

67

GPQA

85.9

SWE

75

Term

64

ARC

39

LiveB

86

Ctx

1M

67
Mistral: Mistral Large 3 2512

Mistral · 9 months ago

Elo

1568

AA

65

GPQA

83.9

SWE

72

Term

60

ARC

36

LiveB

84

Ctx

262K

68
Z.ai: GLM 5V Turbo

Z.ai · 5 months ago

Elo

1447

AA

55

GPQA

77.9

SWE

64

Term

49

ARC

27

LiveB

71

Ctx

203K

69
NVIDIA Nemotron-3.5 Lightning 30B-A3B

NVIDIA ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

70
Qwen: Qwen3.6 Max Preview

Alibaba · 4 months ago

Elo

1485

AA

72

GPQA

77.0

SWE

60

Term

50

ARC

19

LiveB

71

Ctx

262K

71
Google: Gemini 3.1 Flash Lite

Google · 4 months ago

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

1.0M

72
Google: Nano Banana 2 (Gemini 3.1 Flash Image)

Google · 2 months ago

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

131K

73
Gemini Robotics ER 2

Google DeepMind ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

74
Claude 4.5 Opus

Anthropic · 3 months ago

Elo

1458

AA

GPQA

81.9

SWE

Term

ARC

LiveB

Ctx

500K

75
OpenAI: GPT-5.6 Sol Pro

OpenAI · 1 month ago

Elo

1542

AA

69

GPQA

91.0

SWE

80

Term

69

ARC

23

LiveB

86

Ctx

1.1M

76
Llama 4 405B

Meta · 4 months ago

Elo

1389

AA

GPQA

71.2

SWE

Term

ARC

LiveB

Ctx

128K

77
Qwen Drive 1.0

Alibaba Qwen ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

78
Qwen 3.8

Alibaba ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

79
Claude Gen-5

Anthropic ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

80
Claude Mythos

Anthropic ·

Elo

AA

GPQA

SWE

Term

ARC

LiveB

Ctx

Fresh drops

Latest models.

Newest first — including rumored and announced models that aren't yet on the leaderboard.

N=15
  1. OpenAI: GPT-6 Astra· OpenAI

    OpenAI · 2026-09-01 — Path to Astra: critical capabilities and frontier safeguards

    Announcedannounced · 13h ago
  2. Releasedreleased · 1d ago
  3. Releasedreleased · 2d ago
  4. r/singularity · 2026-09-02 · Introducing Gemini 3.8 Flash

    Releasedreleased · 2d ago
  5. Releasedreleased · 3d ago
  6. Releasedreleased · 8d ago
  7. Releasedreleased · 14d ago
  8. OpenRouter · preview endpoint live

    Previewpreview · 14d ago
  9. r/LocalLLaMA · official Z.ai announcement on 2026-08-14

    Releasedreleased · 17d ago
  10. Releasedreleased · 21d ago
  11. Ars Technica · Google announcement on 2026-08-13

    Releasedreleased · 22d ago
  12. Releasedreleased · 23d ago
  13. Releasedreleased · 23d ago
  14. Releasedreleased · 23d ago
  15. TechCrunch · Meta open-weight Muse Glimmer launch

    Releasedreleased · 26d ago

Score Sources

Aggregated by GPT-5.6 Terra

Every score in the table above is aggregated from these public leaderboards. Click any column header to sort the leaderboard; the ↗ icon opens the underlying source. Model names link to their OpenRouter page.

How to read this board

Editorial

The table above is a composite, not a poll. Each model's position comes from reconciling independent public evaluations — human preference votes on LMArena, the independently-run Artificial Analysis index, contamination-resistant LiveBench, graduate-level GPQA Diamond, and execution-graded coding benchmarks like SWE-bench Verified and Terminal-Bench. A model earns the top slot only when it leads across the widest set of those signals at once; winning a single benchmark is never enough.

Treat gaps of a few points as ties. Every evaluation on this page has a noise band of several points, and lab-reported figures regularly differ from independently-run ones — when they disagree, we weight the independent run. Below the top handful of models, scores are best read as estimates with the source attached, not as precise measurements.

"SOTA" is also time-stamped. The current leader, Anthropic: Claude Fable 5.1, has held the position for roughly 4 days since its public release. The frontier typically turns over every few weeks, so the date under the hero matters as much as the name: a SOTA claim from last month is a historical fact, not a current one.

If you are choosing a model rather than watching the race, don't start here — start with the job. The use-case leaderboards for coding, reasoning and agentic work re-weight the same underlying scores for each workload, and the methodology page documents every weight and guardrail.

FAQ

Common questions.

What is the SOTA LLM right now?
The current state-of-the-art LLM on our aggregated leaderboard is shown in the hero at the top of this page. Rankings are refreshed every couple of hours from LMArena, Artificial Analysis, LiveBench, GPQA, SWE-bench, Terminal-Bench and more.
What does SOTA mean in AI?
SOTA stands for state of the art — the model that currently leads on the benchmarks a task cares about. For a general SOTA we aggregate across every major public evaluation; for task-specific SOTA see /coding, /reasoning, or /agentic.
Which LLM is best for coding?
The dedicated coding leaderboard at /coding ranks models by a weighted composite of SWE-bench Verified, Aider Polyglot, Terminal-Bench and LiveBench — the benchmarks that actually predict developer productivity.
How often is this leaderboard updated?
Every ~2 hours. The exact last-refresh timestamp is shown under the H1 and in the footer of every page.

Signal Stream

24 items
  1. r/LocalLLaMA· 2h ago

    Qwen3.8 27b for agentic coding and next .... what?

    First, I'd like to thank the Qwen and Unsloth teams for the Qwen3.8 27b UD Q4_K_XL. Fits the poor 24GB of 3090 VRAM with 100k context at Q8 and works phenomenally well! Imho if theres anything that can threaten Anthropic/OpenAI profits is not another frontier model but actually these small ones you can run fast locally that can do 80..90% of mundane work for hours without paying a single dolla

  2. r/LocalLLaMA· 2h ago

    I've found myself using Local LLM's like 3D printers.

    Anyone who has a 3D printer and get use of it finds it incredibly useful for those odd jobs around the house, a missing bracket, a cable router, steam deck holder and so on. In the past if I was missing an app or useful software, a game I'd do the lazy thing, even though I can and have coded in the past, its "effort" I'll just go and buy or download the latest and greatest. Earli

  3. r/LocalLLaMA· 2h ago

    The OpenAI Huggingface incident from an agents POV

    Full credits to @artificialisabel from X!   submitted by   /u/iPingWine [link]   [comments]

  4. Hacker News· 3h ago

    Could Anthropic have solved Navier–Stokes?

    Article URL: https://twitter.com/ElliotGlazer/status/2096076054133952516 Comments URL: https://news.ycombinator.com/item?id=49573480 Points: 1 # Comments: 0

  5. Hacker News· 3h ago

    Claude Fable 5.1 and Mythos 5.1: The System Card

    Article URL: https://thezvi.substack.com/p/claude-fable-51-and-mythos-51-the Comments URL: https://news.ycombinator.com/item?id=49573407 Points: 1 # Comments: 0

  6. Hacker News· 4h ago

    Ask HN: Replicate ChatGPT/Anthropic Voice Mode?

    Hi, I'm interested in how to have my own phone number where I can call in and get a similar experience to ChatGPT voice mode or Anthropic voice mode. Anyone have experience setting this up? I've got something rudimentary setup with Parakeet for STT, Elevenlabs for TTS, Claude for the brains, and Twilio for the phone number handling. It would have impressed someone in 2024 but... Comments URL: http

  7. r/LocalLLaMA· 5h ago

    How to tune llama.cpp codebase to add custom supported commands for AMD Vulkan?

    Looking for any guidance regarding this, I want to tune my existing llama.cpp codebase so I can potentially implement custom supported commands (not random commands) for my existing amd vulkan setup, I don’t have plans to switch to Linux, I’m using window 11, so ROCm isn’t supported here The reason I’m doing this is because I’m looking forward to optimise the existing configuration, so I can poten

  8. r/LocalLLaMA· 5h ago

    which model is good for detecting deflection?

    I want the answers generated by frontier LLMs or base model LLM answers to be reviewed by some uncensored or abliterated small model. The job is this model (preferably small model) is just to detect deflection in the answers. The problem I am facing is uncensored SLM usually agrees on everything we give input. So the generated answer is also input for it and system prompt is input too.   submi

  9. Hacker News· 5h ago

    Aegis – Inline security sidecar and eBPF sandbox for LLM agents

    Article URL: https://aegiscruc.io Comments URL: https://news.ycombinator.com/item?id=49573010 Points: 1 # Comments: 0

  10. r/LocalLLaMA· 5h ago

    Qwen3.8-27B beat the Wikipedia game in 6 clicks.

    Used qwen3.8-27b in Opencode to make this silly mini-game because I'm not sober: ``` We are going to play a game, it will be the Wikipedia game. The Wikipedia game has the following rules: You will have a Wikipedia article set as a starting point. You will have a Wikipedia article set as an ending point. Your objective is to reach the the end point, which is an article completely separate from

  11. r/LocalLLaMA· 5h ago

    AMD unveils Threadripper Halo Station

    AMD Threadripper Halo Station CPU Ryzen Threadripper PRO 9995WX (Zen 5, "Shimada Peak") 96 cores / 192 threads Up to 5.4 GHz boost 384 MB L3 cache 350 W TDP 8-channel DDR5 128 PCIe 5.0 lanes System Memory 2 TB DDR5 (as shown at IFA) Accelerators 2 x Liquid Cooled AMD Instinct MI350P (CDNA 4) 144 GB HBM3E per card, up to 4 TB/s per card 288 GB total HBM3E Up to 600 W TBP per card PCIe 5.0

  12. Hacker News· 5h ago

    GPT-6 Astra in code review: Gains, privacy, and cost

    Article URL: https://www.coderabbit.ai/blog/gpt-6-astra-code-review-evaluation Comments URL: https://news.ycombinator.com/item?id=49572875 Points: 6 # Comments: 1

  13. r/singularity· 5h ago

    Why would Google or NAVER partner with an independent AI web index?

    Companies such as Parallel and Keenable are building new web indexes designed specifically for AI agents. Google has already integrated Parallel despite owning one of the world’s largest search indexes. Why would Google or NAVER partner with these companies instead of building the technology internally? Does this validate independent AI search as a major new category, or will these startups eventu

  14. r/singularity· 5h ago

    One Shot Galaxy

    Same prompt on Opus 4.8, Fable 5.1, and GPT 6 Astra.   submitted by   /u/LightningMcLovin [link]   [comments]

  15. r/LocalLLaMA· 5h ago

    Chalk one up for the frontier model...

    I just spent the last 2 hours of my life on a Friday night debugging a strange error in a prod CLI app. EF core was receive a readonlyspan during a Contains query. Normally, this query converted to a WHERE [col] IN (...) , but for some reason, after an update, it started choking, despite no code change. The same exact code runs in a separate website docker image fine, no problem. I put Qwen 3.8 27

  16. r/singularity· 6h ago

    Fat Little Boy in an LLM Candy-shop

    That’s me. Lately, I’ve been a fat little boy in the LLM candy-shop. After absolutely gorging myself on the Chinese sweets, K3, GLM 5.3, and Qwen 3.8 max, Elon lured me back to America with Grok 4.6 Panda Express. Then just the other day, boom. Zuck comes out of left field with Cherry Coke aka Muse Spark 1.3. Today, Asstra drops and in my gluttony, I’m using it to orchestrate Muse Spark workers. A

  17. r/singularity· 6h ago

    AA Intelligence Index Changes

    "Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming"   submitted by   /u/poigre [link]   [comments]

  18. r/singularity· 6h ago

    Anthropic Possibly Tackles Its First Millennium Prize Problem

      submitted by   /u/ResultBackground2450 [link]   [comments]

  19. r/LocalLLaMA· 7h ago

    RTX 4090 48GB longevity

    Modified 4090 48GB has been out for a while. I remember a lot of people were buying them at the time. A lot of people were also complaining that they are meant to fail, that they scam etc. I have a few questions to people people who bought these. How is longevity of these cards? Do they still work without issues? Any failure rate? Do they use the same Nvidia drivers that regular 4090 or 4090D uses

  20. Hacker News· 8h ago

    Claude Fable 5.1 vs. GPT-6 Astra, who wins on the 3D modeling?

    Article URL: https://github.com/PhiloLabs/fable51-worlds/tree/main Comments URL: https://news.ycombinator.com/item?id=49572129 Points: 2 # Comments: 2

  21. r/singularity· 8h ago

    AI will turn every profession into an art form instead of something to be productive on

    I feel like a lot of the AI usage is opportunistic, we use it just because we can. Me first, that was my reasoning for using AI in a lot of work. But let's imagine we hit a cap of "infinite productivity", then will it really matter if we produce hand-made or using AI ? All needs everywhere will be fulfilled, for free, in every way possible. We won't feel more competitive by using

  22. r/LocalLLaMA· 8h ago

    Instructions working well for qwen3.8

    Important context: this is about preserve_thinking false stacks and makes no sense if you don't have that working end-to-end with your harness and llamacpp backend already sending reasoning_content and removing it. This is a bit tricky config-wise in llamacpp and your harness and not the default. I'm assuming the reader here already has a lot of prior knowledge. My entire goal is always to

  23. Hacker News· 8h ago

    GPT-6 Astra: First Experience (high)

    Article URL: https://chatgpt.com/share/6a9b671d-b780-83ea-bd5d-b0ea39f7d9b3 Comments URL: https://news.ycombinator.com/item?id=49571948 Points: 2 # Comments: 1

  24. r/singularity· 8h ago

    GPT 6 Astra debuts with a 350 point lead on VoxelBench

      submitted by   /u/LightVelox [link]   [comments]