Cognition(@cognition_labs)

On FrontierCode (Extended), our benchmark for real-world engineering tasks that grades mergeability ...

6.6内容质量

TL;DR · AI 摘要

Cognition on X: "On FrontierCode Extended , our benchmark for real-world engineering tasks that grades mergeability and...

核心要点

  • 主题聚焦:On FrontierCode Extended , our benchmark for rea
  • 来源:Cognition(@cognition_labs),建议结合原文判断细节。
  • AI 分析暂不可用,本条为保底评分与摘要。
#AI#编程
打开原文

Cognition on X: "On FrontierCode (Extended), our benchmark for real-world engineering tasks that grades mergeability and quality, Sonnet 5 scores 53.8% and has a 57.6% pass rate (higher than Opus 4.8). Note: These relative rankings may change slightly with coming adjustments to FrontierCode." / X

Cognition

@cognition

Replying to

On FrontierCode (Extended), our benchmark for real-world engineering tasks that grades mergeability and quality, Sonnet 5 scores 53.8% and has a 57.6% pass rate (higher than Opus 4.8). Note: These relative rankings may change slightly with coming adjustments to FrontierCode.

devin.ai/blog/claude-so…

Claude Sonnet 5 is now available in Devin

From devin.ai

6:21 PM · Jun 30, 2026

4.4K

Views

1

3

31

Read 1 reply