Thomas Wolf(@Thom_Wolf)

You should probably go give @poolsideai a follow on Hugging Face. These folks are on a roll, releas...

8.5内容质量
You should probably go give @poolsideai a follow on Hugging Face.

These folks are on a roll, releas...

TL;DR · AI 摘要

Poolside AI的Laguna S2.1是当前在本地运行的最佳编码模型之一,其在Terminal-Bench 2.1和DeepSWE基准测试中表现优异。

核心要点

  • Laguna S2.1在Terminal-Bench 2.1得分70.2,超越5-25倍参数量的模型
  • Poolside AI每月发布一个代理编码模型,Laguna S2.1支持单DGX Spark/Mac本地运行
  • DeepSWE基准测试显示其在长周期任务中表现突出

结构提纲

按章节快速跳转。

  1. Thomas Wolf推荐关注Poolside AILaguna S2.1编码模型。

  2. Laguna S2.1在Terminal-Bench 2.1得分70.2,超越同规模模型。

  3. DeepSWE长周期任务中表现优于多个大模型。

  4. Poolside AI每月发布一个代理编码模型,持续迭代优化。

思维导图

用一张图看清主题之间的关系。

查看大纲文本(无障碍 / 无 JS 友好)
  • Laguna S2.1模型介绍
    • 性能优势
      • Terminal-Bench 2.1得分70.2
      • 超越5-25倍参数量模型
    • 技术特性
      • 支持本地运行(DGX Spark/Mac)
      • DeepSWE基准测试通过
    • 开发团队
      • Poolside AI每月发布新模型

金句 / Highlights

值得收藏与分享的关键句。

#AI模型#编码模型#Hugging Face#基准测试
打开原文

Thomas Wolf on X: "You should probably go give @poolsideai a follow on Hugging Face. These folks are on a roll, releasing one agentic coding models every month and the latest one Laguna S2.1 is quite possibly the best coding model you can run locally (single DGX spark or Mac) atm https://t.co/74uQy16ksa" / X

Thomas Wolf

@Thom_Wolf

You should probably go give

@

poolsideai

a follow on Hugging Face. These folks are on a roll, releasing one agentic coding models every month and the latest one Laguna S2.1 is quite possibly the best coding model you can run locally (single DGX spark or Mac) atm

Poolside

@poolsideai

Jul 21

Replying to

Laguna S 2.1 is, as far as we can measure, the most capable agentic coding model in its weight class. On Terminal-Bench 2.1 it scores 70.2, sitting beside models 5–25x its size and ahead of several of them. And on DeepSWE from

datacurve

, the hardest long-horizon benchmark we

Show more

8:46 AM · Jul 22, 2026

11.5K

Views

12

0

1

2

17

7

122

16

6