Hesam(@Hesamation)

WAIT… Astra crushes Fable 5.1 on every benchmark but makes 0 progress on the AA Intelligence score? > same 61 as Sol and Grok 4.6 > 51 on Agentic Index vs 59 for Grok > and then it saturates ARC-AGI-3? this is the biggest discrepancy I’ve seen a model’s evals.

7.5内容质量
WAIT… Astra crushes Fable 5.1 on every benchmark but makes 0 progress on the AA Intelligence score?

> same 61 as Sol and Grok 4.6
> 51 on Agentic Index vs 59 for Grok
> and then it saturates ARC-AGI-3?

this is the biggest discrepancy I’ve seen a model’s evals.

TL;DR · AI 摘要

Astra模型在多数基准测试中表现优异,但AA Intelligence得分未提升,且成本较高。

核心要点

  • Astra在Fable 5.1基准中得分与Grok 4.6相同,但AA Intelligence得分未提升。
  • Agentic Index得分51,低于Grok的59,且价格是GPT-5.6的2.5倍。
  • Astra在ARC-AGI-3测试中表现饱和,但未解决AA Intelligence的性能瓶颈。

结构提纲

按章节快速跳转。

  1. 指出Astra模型在多数基准测试中表现优异但存在关键短板。

  2. Astra在Fable 5.1ARC-AGI-3中表现突出,但AA Intelligence得分停滞。

  3. Astra的Agentic Index得分低于Grok,且价格是GPT-5.6的2.5倍。

  4. 模型在不同基准间的性能差异揭示了潜在的优化瓶颈。

思维导图

用一张图看清主题之间的关系。

查看大纲文本(无障碍 / 无 JS 友好)
  • Astra模型评估
    • 基准测试表现
      • Fable 5.1得分61
      • ARC-AGI-3表现饱和
    • 性能短板
      • AA Intelligence得分停滞
      • Agentic Index 51(低于Grok)
    • 成本问题
      • 价格是GPT-5.6的2.5倍

金句 / Highlights

值得收藏与分享的关键句。

#AI模型#基准测试#成本分析#模型对比
打开原文

ℏεsam on X: "WAIT… Astra crushes Fable 5.1 on every benchmark but makes 0 progress on the AA Intelligence score? > same 61 as Sol and Grok 4.6 > 51 on Agentic Index vs 59 for Grok > and then it saturates ARC-AGI-3? this is the biggest discrepancy I’ve seen a model’s evals." / X

ℏεsam

@Hesamation

WAIT… Astra crushes Fable 5.1 on every benchmark but makes 0 progress on the AA Intelligence score? > same 61 as Sol and Grok 4.6 > 51 on Agentic Index vs 59 for Grok > and then it saturates ARC-AGI-3? this is the biggest discrepancy I’ve seen a model’s evals.

Artificial Analysis

@ArtificialAnlys

8h

GPT-6 Astra makes significant gains in the Artificial Analysis Coding Agent Index, scoring equal to Fable 5 at lower cost. In the Intelligence Index, it uses fewer tokens than GPT-5.6 Sol for similar performance, but this is outweighed by higher prices Pricing is 2.5x GPT-5.6

Show more

10:19 PM · Sep 3, 2026

26.8K

Views

18

8

209

25