How to Engineer AI Inference Systems with Philip Kiely - #766

播客收听
问这期播客
会先在本集摘要、章节、转录和笔记里找答案。
TL;DR · AI 摘要
Philip Kiely discusses the critical aspects of AI inference engineering, including its importance, key technologies, and best practices.
核心要点
- Inference is crucial for AI workloads.
- Understanding 'the knobs' improves product design.
- Specialized runtimes enhance performance.
结构提纲
按章节快速跳转。
思维导图
用一张图看清主题之间的关系。
查看大纲文本(无障碍 / 无 JS 友好)
- AI Inference Engineering
金句 / Highlights
值得收藏与分享的关键句。
Inference has become the stickiest and most critical workload in AI.
Understanding ‘the knobs’ of inference lets teams design better products and SLAs.
We trace the inference maturity journey from closed APIs to dedicated deployments and in-house platforms.
章节
- 要点
Inference is crucial for AI workloads.
Inference is crucial for AI workloads.
- 要点
Understanding 'the knobs' improves product design.
Understanding 'the knobs' improves product design.
- 要点
Specialized runtimes enhance performance.
Specialized runtimes enhance performance.
转录
这期还没有可搜索转录。后续抓到带时间戳的内容后会自动补到这里。
节目笔记
How to Engineer AI Inference Systems | TWIML - The Voice of Machine Learning & AI
Twiml icon youtubeTwiml icon X/twitterTwiml icon linkedinTwiml icon FacebookTwiml icon instagram
[](https://twimlai.com/)
How to Engineer AI Inference Systems with Philip Kiely
EPISODE 766
|
APRIL 30, 2026
Watch
Follow








Share




Don't Miss an Episode!_Join our mailing list for episode summaries and other updates._
Don't fill this out if you're human:
First Name
Last Name
JOIN LIST
About this Episode
In this episode, Philip Kiely, head of AI education at Baseten, joins us to unpack the fast-evolving discipline of inference engineering. We explore why inference has become the stickiest and most critical workload in AI, how it blends GPU programming, applied research, and large-scale distributed systems, and where the line sits between inference and model serving. Philip shares how research-to-production can move in hours, not months, and why understanding “the knobs” of inference—batching, quantization, speculation, and KV cache reuse—lets teams design better products and SLAs. We trace the inference maturity journey from closed APIs to dedicated deployments and in-house platforms, discuss GPU lifecycles, and survey today’s runtime landscape, including vLLM, SGLang, and TensorRT LLM. Finally, we look ahead to agents and multimodality, making the case for specialized, workload-specific runtimes when performance and efficiency matter most.
About the Guest
#### Philip Kiely Baseten
Connect with Philip
Resources
- Inference Engineering Book
- Baseten
- PolarQuant: Quantizing KV Caches with Polar Transformation
- NVIDIA TensorRT
- NVIDIA Hopper Architecture
- NVIDIA Blackwell Architecture
- NVIDIA Ampere Architecture
- The path to ubiquitous AI
- Gemini Enterprise Agent Platform
- Amazon Bedrock
- Microsoft Foundry
- Wispr Flow
- Jane Street Signs $6 Billion AI Cloud Agreement with CoreWeave
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
- AWS and Cerebras Collaboration Aims to Set a New Standard for AI Inference Speed and Performance in the Cloud
- NVIDIA H200 GPU
- NVIDIA GB300 NVL72
- NVIDIA H100 GPU
- NVIDIA L4 Tensor Core GPU
- NVIDIA L40 GPU
[](https://twimlai.com/)
© 2026 CloudPulse Strategies
All rights reserved
About TWIML+
Popular Content+
Connect with Us+
Our Policies+
© 2026 CloudPulse Strategies
All rights reserved