Hot take on METR’s new graph that so many people are flipping about today. • Claude Code is a real ...

TL;DR · AI 摘要
Gary Marcus评论METR的新图表,指出其仅展示了50%的成功率,而非100%,强调了GenAI的可靠性问题及软件任务局限性。
核心要点
- METR的新图表仅展示50%成功率
- GenAI的可靠性问题未解决
- 图表仅限于软件任务
结构提纲
按章节快速跳转。
思维导图
用一张图看清主题之间的关系。
查看大纲文本(无障碍 / 无 JS 友好)
- Gary Marcus评论METR新图表
金句 / Highlights
值得收藏与分享的关键句。
If you read the graph carefully, it is about achieving *50%* success. Not 100 or 99 or even 90.
The key problem with GenAI has been reliability; this graph does not address reliable performance. At all.
It certainly doesn’t tell you that *most* (let alone) all things that humans can do in 16 hours can be done in Mythos, let alone reliably
• Claude Code is a real advance; Mythos probably builds on some of what is learned there. But…
• If you read the graph carefully, it is about achieving *50%* success. Not 100 or 99 or even 90. The" / X
Gary Marcus on X: "Hot take on METR’s new graph that so many people are flipping about today. •Claude Code is a real advance; Mythos probably builds on some of what is learned there. But… •If you read the graph carefully, it is about achieving *50%* success. Not 100 or 99 or even 90. The" / X
Don’t miss what’s happening

Hot take on METR’s new graph that so many people are flipping about today. •Claude Code is a real advance; Mythos probably builds on some of what is learned there. But… •If you read the graph carefully, it is about achieving *50%* success. Not 100 or 99 or even 90. The key problem with GenAI has been reliability; this graph does not address reliable performance. At all. •If you read carefully, it is only about software tasks. Not general intelligence. •It certainly doesn’t tell you that *most* (let alone) all things that humans can do in 16 hours can be done in Mythos, let alone reliably •Aside from this, the graph doesn’t show you *how* the improvements have been made. As noted in my newsletter a lot of the advance in recent months is likely from the incorporation of symbolic tools (like code interpreters, verification, and harnesses) rather than from model scaling per se. As such this a vindication of neurosymbolic AI – but not a proof that LLMs themselves can be perpetually scaled. As such it’s not a proof that another trillion dollars will continue the graph. • Per
, Mythos is not actually off trend on the ECI benchmark, which is a broader measure.
Quote

METR
@METR_Evals
·
18h
We evaluated an early version of Claude Mythos Preview for risk assessment during a limited window in March 2026. We estimated a 50%-time-horizon of at least 16hrs (95% CI 8.5hrs to 55hrs) on our task suite, at the upper end of what we can measure without new tasks.
·
25
13
79
21
Read 25 replies