Inside distributed inference with llm-d ft. Carlos Costa episode artwork

EPISODE · Aug 6, 2025 · 26 MIN

Inside distributed inference with llm-d ft. Carlos Costa

from Technically Speaking with Chris Wright · host Carlos Costa, Chris Wright

Scaling LLM inference for production isn't just about adding more machines, it demands new intelligence in the infrastructure itself. In this episode, we're joined by Carlos Costa, Distinguished Engineer at IBM Research, a leader in large-scale compute and a key figure in the llm-d project. We discuss how to move beyond single-server deployments and build the intelligent, AI-aware infrastructure needed to manage complex workloads efficiently. Carlos Costa shares insights from his deep background in HPC and distributed systems, including: • The evolution from traditional HPC and large-scale training to the unique challenges of distributed inference for massive models. • The origin story of the llm-d project, a collaborative, open-source effort to create a much-needed ""common AI stack"" and control plane for the entire community. • How llm-d extends Kubernetes with the specialization required for AI, enabling state-aware scheduling that standard Kubernetes wasn't designed for. • Key architectural innovations like the disaggregation of prefill and decode stages and support for wide parallelism to efficiently run complex Mixture of Experts (MOE) models. Tune in to discover how this collaborative, open-source approach is building the standardized, AI-aware infrastructure necessary to make massive AI models practical, efficient, and accessible for everyone.

Episode metadata supplied by the publisher feed · Published Aug 6, 2025

Embed this episode

Ready to play

Inside distributed inference with llm-d ft. Carlos Costa

0:00 26:23

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Technically Speaking with Chris Wright?

This episode is 26 minutes long.

When was this Technically Speaking with Chris Wright episode published?

This episode was published on August 6, 2025.

Can I download this Technically Speaking with Chris Wright episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!