lmarena.ai(@lmarena_ai)
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a globa...
8.5内容质量

TL;DR · AI 摘要
Agent Arena平台通过真实任务评估模型性能,Inkling-Small在性价比上表现突出。
核心要点
- Inkling-Small参数量276B,价格仅为Inkling的1/2.2
- Agent Arena使用因果追踪方法评估模型长期任务表现
- Thinking Machines成为美国最便宜的顶级模型提供商
结构提纲
按章节快速跳转。
- §平台介绍
Agent Arena通过全球用户任务评估模型实际表现
- ·模型表现
Inkling-Small以276B参数量位列开放模型第12
- ›价格对比
Inkling-Small单任务成本0.20美元,优于Inkling的0.09美元
- ·评估方法
采用因果追踪方法衡量模型长期任务完成效果
- ›技术指标
Inkling-Small拥有12B活跃参数,参数量仅为Inkling的1/4
思维导图
用一张图看清主题之间的关系。
查看大纲文本(无障碍 / 无 JS 友好)
- Agent Arena评估体系
- 模型评估方法
- 因果追踪评估
- 核心模型
- Inkling-Small
- Inkling
- 性能指标
- 参数量276B
- 价格0.09美元/任务
金句 / Highlights
值得收藏与分享的关键句。
Inkling-Small以276B参数量实现与Inkling相当的性能,价格仅为1/2.2
Agent Arena通过因果追踪方法评估模型在长期任务中的实际表现
Thinking Machines成为美国最便宜的顶级模型提供商,单任务成本0.09美元
#Agent Arena#模型评估#Inkling-Small#Thinking Machines
打开原文Arena.ai on X: "In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to" / X
[](https://x.com/)
Arena.ai on X: "In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to"
-  Arena.ai @arena 11h Inkling-Small by @thinkymachines has hit the Agent Arena. Among open models, it is #12 with a net-improvement score of -7.0%. It delivers similar results with Inkling at less than half the price per task ($0.20 vs. $0.09). Inkling and Inkling-Small make @thinkymachines the top US Show more  [](https://x.com/arena/status/2089826844854112371/photo/1)  [](https://x.com/arena/status/2089826844854112371/photo/2) Thinking Machines @thinkymachines Jul 30 Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available. thinkingmachines.ai/news/inkling-s… Fine-tune it on Tinker today, or chat with Show more 10 11 124 21K
-  Arena.ai @arena In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. See the full Agent Arena leaderboard: arena.ai/leaderboard/ag… From arena.ai 9:28 PM · Aug 18, 20263K Views 1 2
Log in or sign up for X
See what’s happening and join the conversation
Continue with phoneContinue with Apple Continue with Google
or