TWIML AI Podcast播客53:19

How to Find the Agent Failures Your Evals Miss with Scott Clark - #767

7.5内容质量
How to Find the Agent Failures Your Evals Miss with Scott Clark - #767

播客收听

时长 53:19原播客页面

问这期播客

会先在本集摘要、章节、转录和笔记里找答案。

TL;DR · AI 摘要

Scott Clark讨论了如何通过层次化的可观测性方法发现复杂LLM系统中的未检测到的代理失败,强调了在线分析和自适应方法的重要性。

核心要点

  • 使用层次化的可观测性方法发现未知问题
  • 在线分析和自适应方法对非稳定模型至关重要
  • OpenTelemetry和GenAI语义规范有助于系统监控

结构提纲

按章节快速跳转。

  1. Scott Clark讨论了如何通过层次化的可观测性方法发现复杂LLM系统中的未检测到的代理失败。

  2. 包括日志记录、已知信号监控和在线分析,以发现未知问题。

  3. 用于收集系统的运行时数据。

  4. 用于监测已知的问题信号。

  5. 用于发现未知的未知问题。

  6. 对于非稳定模型至关重要,能够生成评估、护栏和训练数据。

思维导图

用一张图看清主题之间的关系。

查看大纲文本(无障碍 / 无 JS 友好)
  • 发现LLM系统中的代理失败

金句 / Highlights

值得收藏与分享的关键句。

章节

  1. 要点

    使用层次化的可观测性方法发现未知问题

    使用层次化的可观测性方法发现未知问题

  2. 要点

    在线分析和自适应方法对非稳定模型至关重要

    在线分析和自适应方法对非稳定模型至关重要

  3. 要点

    OpenTelemetry和GenAI语义规范有助于系统监控

    OpenTelemetry和GenAI语义规范有助于系统监控

转录

这期还没有可搜索转录。后续抓到带时间戳的内容后会自动补到这里。

#LLM#可观测性#机器学习#代理失败

节目笔记

How to Find the Agent Failures Your Evals Miss | TWIML - The Voice of Machine Learning & AI

AboutContactNewsletter

Twiml icon youtubeTwiml icon X/twitterTwiml icon linkedinTwiml icon FacebookTwiml icon instagram

[](https://twimlai.com/)

Twiml icon youtube

Twiml icon linkedin

Twiml icon X/twitter

Twiml icon Facebook

Twiml icon instagram

How to Find the Agent Failures Your Evals Miss with Scott Clark

EPISODE 767

|

MAY 7, 2026

Watch

Play

Follow

![Image 1Apple Podcasts](https://podcasts.apple.com/us/podcast/the-twiml-ai-podcast-formerly-this-week-in-machine/id1116303051)

![Image 2Spotify](https://open.spotify.com/show/2sp5EL7s7EqxttxwwoJ3i7)

![Image 3YouTube](https://www.youtube.com/twimlai)

![Image 4Overcast](https://overcast.fm/itunes1116303051)

![Image 5Podcast Addict](https://podcastaddict.com/podcast/2990273)

![Image 6Castbox](https://castbox.fm/channel/id4477675)

![Image 7Pocket Casts](https://pocketcasts.com/podcast/the-twiml-ai-podcast-formerly-this-week-in-machine-learning-artificial-intelligence/8fc91c30-03a7-0134-9c92-59d98c6b72b8)

![Image 8RSS](https://feeds.megaphone.fm/MLN2155636147)

Share

![Image 9LinkedIn](https://twimlai.com/podcast/twimlai/how-find-agent-failures-your-evals-miss)

![Image 10X](https://twimlai.com/podcast/twimlai/how-find-agent-failures-your-evals-miss)

![Image 11Facebook](https://twimlai.com/podcast/twimlai/how-find-agent-failures-your-evals-miss)

![Image 12Reddit](https://twimlai.com/podcast/twimlai/how-find-agent-failures-your-evals-miss)

Don't Miss an Episode!_Join our mailing list for episode summaries and other updates._

Don't fill this out if you're human:

First Name

Last Name

Email

JOIN LIST

About this Episode

In this episode, Scott Clark, co-founder and CEO of Distributional, joins us to explore how teams can reliably operate and improve complex LLM systems and agents in production. Scott introduces a Maslow’s hierarchy of observability: telemetry for logging, monitoring for known signals, and post-production or online analytics to surface unknown unknowns. We dig into examples of real-world failures Scott’s team has seen in production systems, such as “lazy” tool-use hallucinations that standard evals miss, and how mapping traces into vector fingerprints enables clustering and topic discovery to uncover emergent behaviors. Scott explains how analytics can feed the data flywheel by generating evals, guardrails, and training data, and why online, adaptive approaches are essential for non-stationary models. We also touch on practical how-to’s such as instrumentation with OpenTelemetry, the GenAI semantic conventions, and the role of dedicated analytics tools.

About the Guest

#### Scott Clark Distributional

Connect with Scott

Image 13
Image 13
Image 14
Image 14
Image 15
Image 15
Image 16
Image 16
Image 17
Image 17
Image 18
Image 18

Thanks to our sponsor Distributional

This show is brought to you by our friends at Distributional, the AI analytics platform built for teams that are serious about agent quality. Distributional finds patterns in production agent traces, creating actionable insights with suggestions for new evals, refined guardrails, and improvements to your agent based on real usage. Don't take their word for it, use Distributional yourself for free. Go to app.dbnl.com to create your free hosted account that also comes with a free LLM endpoint to power evals and analytics. Or install a self-hosted version of Distributional for free in your own environment. To learn more, visit dbnl.com and start improving production agent quality today.

Image 19: Distributional Logo
Image 19: Distributional Logo

Resources

Related Topics

![Image 20: Cover: TWIML Presents: AI Agents AI Agents](https://twimlai.com/podcast/twimlai/topics/ai-agents)

Related Episodes

Image 21: Cover Image: Rashmi Shetty - Podcast Interview
Image 21: Cover Image: Rashmi Shetty - Podcast Interview

765

How Capital One Delivers Multi-Agent Systems

Image 22: Cover Image: Siddhant Pardeshi - Podcast Interview
Image 22: Cover Image: Siddhant Pardeshi - Podcast Interview

763

Agent Swarms and Knowledge Graphs for Autonomous Software Development

Image 23: Cover Image: Aakanksha Chowdhery - Podcast Interview
Image 23: Cover Image: Aakanksha Chowdhery - Podcast Interview

759

Rethinking Pre-Training for Agentic AI

Image 24: Cover Image: Zain Asgar - Podcast Interview
Image 24: Cover Image: Zain Asgar - Podcast Interview

757

Scaling Agentic Inference Across Heterogeneous Compute

Image 25: Cover Image: Devi Parikh - Podcast Interview
Image 25: Cover Image: Devi Parikh - Podcast Interview

756

Proactive Agents for the Web

Image 26: Cover Image: Filip Kozera - Podcast Interview
Image 26: Cover Image: Filip Kozera - Podcast Interview

741

Context Engineering for Productive AI Agents

Image 27: Cover Image: Kwindla Kramer - Podcast Interview
Image 27: Cover Image: Kwindla Kramer - Podcast Interview

739

Building Voice AI Agents That Don’t Suck

Image 28: Cover Image: Vijoy Pandey - Podcast Interview
Image 28: Cover Image: Vijoy Pandey - Podcast Interview

737

Building the Internet of Agents

Image 29: Cover Image: Mahesh Sathiamoorthy - Podcast Interview
Image 29: Cover Image: Mahesh Sathiamoorthy - Podcast Interview

731

From Prompts to Policies: How RL Builds Better AI Agents

Image 30: Cover Image: Josh Tobin - Podcast Interview
Image 30: Cover Image: Josh Tobin - Podcast Interview

730

How OpenAI Builds AI Agents That Think and Act

Image 31: Cover Image: Victor Dibia - Podcast Interview
Image 31: Cover Image: Victor Dibia - Podcast Interview

718

AI Trends 2025: AI Agents and Multi-Agent Systems

Image 32: Cover Image: Abhijit Bose - Podcast Interview
Image 32: Cover Image: Abhijit Bose - Podcast Interview

714

Evolving MLOps Platforms for Generative AI and Agents

Image 33: Cover Image: Daniel Jeffries - Podcast Interview
Image 33: Cover Image: Daniel Jeffries - Podcast Interview

713

Why Agents Are Stupid & What We Can Do About It

Image 34: Cover Image: Sunil Mallya - Podcast Interview
Image 34: Cover Image: Sunil Mallya - Podcast Interview

708

An Agentic Mixture of Experts for DevOps

Image 35: Cover Image: Scott Stephenson - Podcast Interview
Image 35: Cover Image: Scott Stephenson - Podcast Interview

707

Building AI Voice Agents

Image 36: Cover Image: Arvind Narayanan - Podcast Interview
Image 36: Cover Image: Arvind Narayanan - Podcast Interview

704

AI Agents: Substance or Snake Oil

Image 37: Cover Image: Shreya Shankar - Podcast Interview
Image 37: Cover Image: Shreya Shankar - Podcast Interview

703

AI Agents for Data Analysis

Image 38: Cover Image: Shengran Hu - Podcast Interview
Image 38: Cover Image: Shengran Hu - Podcast Interview

700

Automated Design of Agentic Systems

Image 39: Cover Image: Harrison Chase - Podcast Interview
Image 39: Cover Image: Harrison Chase - Podcast Interview

698

The Building Blocks of Agentic Systems

Image 40: Cover Image: Joon Park - Podcast Interview
Image 40: Cover Image: Joon Park - Podcast Interview

632

Modeling Human Behavior with Generative Agents

[](https://twimlai.com/)

© 2026 CloudPulse Strategies

All rights reserved

About TWIML+

Popular Content+

Connect with Us+

Our Policies+

© 2026 CloudPulse Strategies

All rights reserved