TWIML AI Podcast播客54:51

How to Engineer AI Inference Systems with Philip Kiely - #766

8.5内容质量
How to Engineer AI Inference Systems with Philip Kiely - #766

播客收听

时长 54:51原播客页面

问这期播客

会先在本集摘要、章节、转录和笔记里找答案。

TL;DR · AI 摘要

Philip Kiely discusses the critical aspects of AI inference engineering, including its importance, key technologies, and best practices.

核心要点

  • Inference is crucial for AI workloads.
  • Understanding 'the knobs' improves product design.
  • Specialized runtimes enhance performance.

结构提纲

按章节快速跳转。

  1. Philip Kiely讨论了AI推理工程的关键方面。

  2. Philip Kiely是Baseten的AI教育负责人。

  3. AI推理成为AI工作负载中最关键的部分。

  4. GPU编程、应用研究和大规模分布式系统是AI推理的关键技术。

  5. AI推理成熟度从闭源API发展到专用部署。

  6. 未来将关注代理和多模态,强调性能和效率。

思维导图

用一张图看清主题之间的关系。

查看大纲文本(无障碍 / 无 JS 友好)
  • AI Inference Engineering

金句 / Highlights

值得收藏与分享的关键句。

章节

  1. 要点

    Inference is crucial for AI workloads.

    Inference is crucial for AI workloads.

  2. 要点

    Understanding 'the knobs' improves product design.

    Understanding 'the knobs' improves product design.

  3. 要点

    Specialized runtimes enhance performance.

    Specialized runtimes enhance performance.

转录

这期还没有可搜索转录。后续抓到带时间戳的内容后会自动补到这里。

#AI#Inference#Engineering

节目笔记

How to Engineer AI Inference Systems | TWIML - The Voice of Machine Learning & AI

AboutContactNewsletter

Twiml icon youtubeTwiml icon X/twitterTwiml icon linkedinTwiml icon FacebookTwiml icon instagram

[](https://twimlai.com/)

Twiml icon youtube

Twiml icon linkedin

Twiml icon X/twitter

Twiml icon Facebook

Twiml icon instagram

How to Engineer AI Inference Systems with Philip Kiely

EPISODE 766

|

APRIL 30, 2026

Watch

Play

Follow

![Image 6Apple Podcasts](https://podcasts.apple.com/us/podcast/the-twiml-ai-podcast-formerly-this-week-in-machine/id1116303051)

![Image 7Spotify](https://open.spotify.com/show/2sp5EL7s7EqxttxwwoJ3i7)

![Image 8YouTube](https://www.youtube.com/twimlai)

![Image 9Overcast](https://overcast.fm/itunes1116303051)

![Image 10Podcast Addict](https://podcastaddict.com/podcast/2990273)

![Image 11Castbox](https://castbox.fm/channel/id4477675)

![Image 12Pocket Casts](https://pocketcasts.com/podcast/the-twiml-ai-podcast-formerly-this-week-in-machine-learning-artificial-intelligence/8fc91c30-03a7-0134-9c92-59d98c6b72b8)

![Image 13RSS](https://feeds.megaphone.fm/MLN2155636147)

Share

![Image 14LinkedIn](https://twimlai.com/podcast/twimlai/how-engineer-ai-inference-systems)

![Image 15X](https://twimlai.com/podcast/twimlai/how-engineer-ai-inference-systems)

![Image 16Facebook](https://twimlai.com/podcast/twimlai/how-engineer-ai-inference-systems)

![Image 17Reddit](https://twimlai.com/podcast/twimlai/how-engineer-ai-inference-systems)

Don't Miss an Episode!_Join our mailing list for episode summaries and other updates._

Don't fill this out if you're human:

First Name

Last Name

Email

JOIN LIST

About this Episode

In this episode, Philip Kiely, head of AI education at Baseten, joins us to unpack the fast-evolving discipline of inference engineering. We explore why inference has become the stickiest and most critical workload in AI, how it blends GPU programming, applied research, and large-scale distributed systems, and where the line sits between inference and model serving. Philip shares how research-to-production can move in hours, not months, and why understanding “the knobs” of inference—batching, quantization, speculation, and KV cache reuse—lets teams design better products and SLAs. We trace the inference maturity journey from closed APIs to dedicated deployments and in-house platforms, discuss GPU lifecycles, and survey today’s runtime landscape, including vLLM, SGLang, and TensorRT LLM. Finally, we look ahead to agents and multimodality, making the case for specialized, workload-specific runtimes when performance and efficiency matter most.

About the Guest

#### Philip Kiely Baseten

Connect with Philip

Image 18
Image 18
Image 19
Image 19
Image 20
Image 20
Image 21
Image 21
Image 22
Image 22

Resources

[](https://twimlai.com/)

© 2026 CloudPulse Strategies

All rights reserved

About TWIML+

Popular Content+

Connect with Us+

Our Policies+

© 2026 CloudPulse Strategies

All rights reserved