Hesam(@Hesamation)
WAIT… Astra crushes Fable 5.1 on every benchmark but makes 0 progress on the AA Intelligence score? > same 61 as Sol and Grok 4.6 > 51 on Agentic Index vs 59 for Grok > and then it saturates ARC-AGI-3? this is the biggest discrepancy I’ve seen a model’s evals.
7.5内容质量

TL;DR · AI 摘要
Astra模型在多数基准测试中表现优异,但AA Intelligence得分未提升,且成本较高。
核心要点
- Astra在Fable 5.1基准中得分与Grok 4.6相同,但AA Intelligence得分未提升。
- Agentic Index得分51,低于Grok的59,且价格是GPT-5.6的2.5倍。
- Astra在ARC-AGI-3测试中表现饱和,但未解决AA Intelligence的性能瓶颈。
结构提纲
按章节快速跳转。
思维导图
用一张图看清主题之间的关系。
查看大纲文本(无障碍 / 无 JS 友好)
- Astra模型评估
- 基准测试表现
- Fable 5.1得分61
- ARC-AGI-3表现饱和
- 性能短板
- AA Intelligence得分停滞
- Agentic Index 51(低于Grok)
- 成本问题
- 价格是GPT-5.6的2.5倍
金句 / Highlights
值得收藏与分享的关键句。
Astra在Fable 5.1中得分61,与Grok 4.6相同但AA Intelligence得分未提升。
Agentic Index得分51(Astra) vs 59(Grok),价格是GPT-5.6的2.5倍。
Astra在ARC-AGI-3测试中表现饱和,但未解决AA Intelligence的性能瓶颈。
#AI模型#基准测试#成本分析#模型对比
打开原文ℏεsam on X: "WAIT… Astra crushes Fable 5.1 on every benchmark but makes 0 progress on the AA Intelligence score? > same 61 as Sol and Grok 4.6 > 51 on Agentic Index vs 59 for Grok > and then it saturates ARC-AGI-3? this is the biggest discrepancy I’ve seen a model’s evals." / X
ℏεsam
@Hesamation
WAIT… Astra crushes Fable 5.1 on every benchmark but makes 0 progress on the AA Intelligence score? > same 61 as Sol and Grok 4.6 > 51 on Agentic Index vs 59 for Grok > and then it saturates ARC-AGI-3? this is the biggest discrepancy I’ve seen a model’s evals.
@ArtificialAnlys
8h
GPT-6 Astra makes significant gains in the Artificial Analysis Coding Agent Index, scoring equal to Fable 5 at lower cost. In the Intelligence Index, it uses fewer tokens than GPT-5.6 Sol for similar performance, but this is outweighed by higher prices Pricing is 2.5x GPT-5.6
Show more
10:19 PM · Sep 3, 2026
26.8K
Views
18
8
209
25