WAIT… Astra crushes Fable 5.1 on every benchmark but makes 0 progress on the AA Intelligence score? > same 61 as Sol and Grok 4.6 > 51 on Agentic Index vs 59 for Grok > and then it saturates ARC-AGI-3? this is the biggest discrepancy I’ve seen a model’s evals.
Hesam(@Hesamation)182 字 (约 1 分钟)
75
Astra模型在多数基准测试中表现优异,但AA Intelligence得分未提升,且成本较高。
入选理由:Astra在Fable 5.1基准中得分与Grok 4.6相同,但AA Intelligence得分未提升。
FeaturedTweet#AI模型#基准测试#成本分析#模型对比中英混合
