How to Find the Agent Failures Your Evals Miss with Scott Clark - #767

播客收听
问这期播客
会先在本集摘要、章节、转录和笔记里找答案。
TL;DR · AI 摘要
Scott Clark讨论了如何通过层次化的可观测性方法发现复杂LLM系统中的未检测到的代理失败,强调了在线分析和自适应方法的重要性。
核心要点
- 使用层次化的可观测性方法发现未知问题
- 在线分析和自适应方法对非稳定模型至关重要
- OpenTelemetry和GenAI语义规范有助于系统监控
结构提纲
按章节快速跳转。
思维导图
用一张图看清主题之间的关系。
查看大纲文本(无障碍 / 无 JS 友好)
- 发现LLM系统中的代理失败
金句 / Highlights
值得收藏与分享的关键句。
Scott介绍了一种Maslow的层次化的可观测性方法,包括日志记录、监控和在线分析。
在线分析能够发现标准评估遗漏的代理失败,如'懒惰'工具使用的幻觉。
在线分析能够生成评估、护栏和训练数据,从而形成数据飞轮效应。
章节
- 要点
使用层次化的可观测性方法发现未知问题
使用层次化的可观测性方法发现未知问题
- 要点
在线分析和自适应方法对非稳定模型至关重要
在线分析和自适应方法对非稳定模型至关重要
- 要点
OpenTelemetry和GenAI语义规范有助于系统监控
OpenTelemetry和GenAI语义规范有助于系统监控
转录
这期还没有可搜索转录。后续抓到带时间戳的内容后会自动补到这里。
节目笔记
How to Find the Agent Failures Your Evals Miss | TWIML - The Voice of Machine Learning & AI
Twiml icon youtubeTwiml icon X/twitterTwiml icon linkedinTwiml icon FacebookTwiml icon instagram
[](https://twimlai.com/)
How to Find the Agent Failures Your Evals Miss with Scott Clark
EPISODE 767
|
MAY 7, 2026
Watch
Follow








Share




Don't Miss an Episode!_Join our mailing list for episode summaries and other updates._
Don't fill this out if you're human:
First Name
Last Name
JOIN LIST
About this Episode
In this episode, Scott Clark, co-founder and CEO of Distributional, joins us to explore how teams can reliably operate and improve complex LLM systems and agents in production. Scott introduces a Maslow’s hierarchy of observability: telemetry for logging, monitoring for known signals, and post-production or online analytics to surface unknown unknowns. We dig into examples of real-world failures Scott’s team has seen in production systems, such as “lazy” tool-use hallucinations that standard evals miss, and how mapping traces into vector fingerprints enables clustering and topic discovery to uncover emergent behaviors. Scott explains how analytics can feed the data flywheel by generating evals, guardrails, and training data, and why online, adaptive approaches are essential for non-stationary models. We also touch on practical how-to’s such as instrumentation with OpenTelemetry, the GenAI semantic conventions, and the role of dedicated analytics tools.
About the Guest
#### Scott Clark Distributional
Connect with Scott
Thanks to our sponsor Distributional
This show is brought to you by our friends at Distributional, the AI analytics platform built for teams that are serious about agent quality. Distributional finds patterns in production agent traces, creating actionable insights with suggestions for new evals, refined guardrails, and improvements to your agent based on real usage. Don't take their word for it, use Distributional yourself for free. Go to app.dbnl.com to create your free hosted account that also comes with a free LLM endpoint to power evals and analytics. Or install a self-hosted version of Distributional for free in your own environment. To learn more, visit dbnl.com and start improving production agent quality today.

Resources
- Distributional
- Distributional App
- Distributional Docs
- Clio: Privacy-Preserving Insights into Real-World AI Use
- Where the goblins came from
- An update on recent Claude Code quality reports
- Datadog
- Statsig
- Braintrust
- Mixpanel
- Decagon
- Harvey AI
- Supporting Rapid Model Development at Two Sigma with Scott Clark & Matthew Adereth - #273
- Bayesian Optimization for Hyperparameter Tuning with Scott Clark - #50
- Democast: Automated Model Tuning with Scott Clark
- Building Real-World LLM Products with Fine-Tuning and More with Hamel Husain - #694
Related Topics

Related Episodes

765
How Capital One Delivers Multi-Agent Systems

763
Agent Swarms and Knowledge Graphs for Autonomous Software Development

759
Rethinking Pre-Training for Agentic AI

757
Scaling Agentic Inference Across Heterogeneous Compute

756

741
Context Engineering for Productive AI Agents

739
Building Voice AI Agents That Don’t Suck

737
Building the Internet of Agents

731
From Prompts to Policies: How RL Builds Better AI Agents

730
How OpenAI Builds AI Agents That Think and Act

718
AI Trends 2025: AI Agents and Multi-Agent Systems

714
Evolving MLOps Platforms for Generative AI and Agents

713
Why Agents Are Stupid & What We Can Do About It

708
An Agentic Mixture of Experts for DevOps

707

704
AI Agents: Substance or Snake Oil

703

700
Automated Design of Agentic Systems

698
The Building Blocks of Agentic Systems

632
Modeling Human Behavior with Generative Agents
[](https://twimlai.com/)
© 2026 CloudPulse Strategies
All rights reserved
About TWIML+
Popular Content+
Connect with Us+
Our Policies+
© 2026 CloudPulse Strategies
All rights reserved