Lenny's Newsletter

Sonnet 5 review: I ran 64 generations to find out if it's worth it

6.9内容质量
Sonnet 5 review: I ran 64 generations to find out if it's worth it

TL;DR · AI 摘要

Sonnet 5 review: I ran 64 generations to find out if it's worth it Playback speed × Share post Share post at current tim...

核心要点

  • 主题聚焦:Sonnet 5 review: I ran 64 generations to find ou
  • 来源:Lenny's Newsletter,建议结合原文判断细节。
  • AI 分析暂不可用,本条为保底评分与摘要。
#AI#编程#产品
打开原文

Sonnet 5 review: I ran 64 generations to find out if it's worth it

Playback speed

×

Share post

Share post at current time

Share from 0:00

0:00

/

Generate transcript

A transcript unlocks clips, previews, and editing.

Sonnet 5 review: I ran 64 generations to find out if it's worth it

🎙 I built the How I AI Bench live using Claude Code, ran 5 frontier models through 64 blind prototype generations, PRDs, and agent voice tests, and the results surprised even me

Claire Vo

Jun 30, 2026

Transcript

I’ve been testing every major frontier model release since the start of the year, and when Anthropic dropped Sonnet 5, I wanted more than a vibe check. I got tired of one-off tests I couldn’t repeat or compare over time, so I built something better: the How I AI Bench, a repeatable eval harness I constructed live using Claude Code while recording this episode. I ran Sonnet 5 blind against four other frontier models (Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro) across PRD quality, prototype generation, agentic task completion, and agent personality. The results were not what I expected.

Listen or watch on YouTube , Spotify , or Apple Podcasts

What you’ll learn:

  • What Anthropic claims Sonnet 5 improves over Sonnet 4.6, and where the benchmark data actually backs that up
  • How I built the How I AI Bench in under 45 minutes using Claude Code, starting from my own stored session history
  • Why I combined human vibe scoring (70%) with LLM as judge scoring (30%) instead of trusting either alone
  • How to set up a local HTML scoring page so you can rate AI outputs on gut feel and export those scores as JSON
  • Which model I recommend for PRDs, which for complex prototypes, and which for chatting with an agent daily

Brought to you by:

Runway —The creative AI platform for images, video and more

Hyperagent —Deploy fleets of agents that handle real work

In this episode, we cover:

( 00:00 ) Sonnet 5 is out

( 01:55 ) What Anthropic claims

( 04:02 ) Why I’m done with one-off vibe checks

( 05:05 ) Building the How I AI Bench live with Claude Code

( 07:42 ) The scoring system

( 10:43 ) Agent voice eval

( 11:57 ) Quick recap

( 13:58 ) Results: The How I AI index leaderboard

( 21:21 ) What I’m improving for the next run

( 22:16 ) Generating a Claire-weighted index

( 23:53 ) Model-by-task recommendations

Tools referenced:

• Claude Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5

• Claude Opus 4.8: https://www.anthropic.com/news/claude-opus-4-8

• GPT-5.5 (OpenAI): https://openai.com/index/introducing-gpt-5-5/

• Gemini 3 Pro (Google DeepMind): https://deepmind.google/models/gemini/pro/

• Cursor: https://www.cursor.com/

Other references:

• SWE-bench Pro (agentic coding benchmark referenced): https://www.swebench.com/

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

Production and marketing by https://penname.co/ . For inquiries about sponsoring the podcast, email [email protected] .

#### Discussion about this video

Comments

Restacks

How I AI

How I AI, hosted by Claire Vo, is for anyone wondering how to actually use these magical new tools to improve the quality and efficiency of their work. In each episode, guests will share a specific, practical, and impactful way they’ve learned to use AI in their work or life. Expect 30-minute episodes, live screen sharing, and tips/tricks/workflows you can copy immediately. If you want to demystify AI and learn the skills you need to thrive in this new world, this podcast is for you.

Listen on

Substack App

Apple Podcasts

Spotify

YouTube

Overcast

Pocket Casts

RSS Feed

Appears in episode

Writes

Claire’s Substack

Recent Episodes

No Figma. No Jira. No docs. How Gusto built a new product line with Claude Code | Eddie Kim (CTO)

Jun 29

GLM 5.2: why I’m replacing Opus in Claude Code with this new model

Jun 24

How Claude Mythos found a 15-year-old bug in Mozilla Firefox | Brian Grinstead

Jun 22

How to design AI agent loops: schedules, goals, and subagents in Claude Code and Codex

Jun 17

How Braintrust uses AI agents, evals, and CI to ship better software | Ankur Goyal

Jun 15

Claude Fable 5 review: what the new Mythos model gets right (and very wrong)

Jun 9

Shopping with Claude: How to find quality brands, automate returns, and buy things that last 100 years | Nicole Ruiz

Jun 8