Inside AMD’s AI Strategy From Edge To Data Center episode artwork

EPISODE · Dec 15, 2025 · 30 MIN

Inside AMD’s AI Strategy From Edge To Data Center

from What's Up with Tech? · host Evan Kirstel

Interested in being a guest? Email us at [email protected] leaps in AI rarely come from one breakthrough. They emerge when hardware design, open software, and real workloads click into place. That’s the story we unpack with AMD’s Ramine Roane: how an open, developer-first approach combined with high-bandwidth memory, chiplet packaging, and a re-architected software stack is reshaping performance and cost from the edge to the largest data centers.We walk through why memory capacity and bandwidth dominate large language model performance, and how MI300X’s 192 GB HBM and advanced packaging unlock bigger contexts and faster token throughput. Ramin explains how Rocm 7 was rebuilt to be modular, smaller to install, and enterprise-ready—so teams can go from single-node experiments to fully orchestrated clusters using Kubernetes, Slurm, and familiar open tools. The highlight: disaggregated and distributed inference. By splitting prefill from decode and adopting expert parallelism, organizations are slashing cost per token by 10–30x, depending on model and topology.The conversation ranges from startup-friendly workflows to hyperscaler deployments, with practical insight into VLLM, SGLang, and why open source now outpaces closed stacks. We also look ahead at where inference runs: the edge is rising. With performance per watt doubling on a steady cadence, AI PCs, laptops, and phones will take on more of the work, enabling privacy, responsiveness, and lower costs. Ramin shares a sober view on quantum computing timelines and a bullish take on the broader compute shift—moving once-sequential problems into massively parallel deep learning that changes what’s even possible.If you care about real performance, total cost of ownership, and developer velocity, this conversation brings a grounded blueprint: open ecosystems, smarter packaging, and inference architectures built for high utilization. Subscribe, share with a colleague who cares about LLM throughput and cost, and leave a quick review to help others find the show.Everyday AI: Your daily guide to grown with Generative AICan't keep up with AI? We've got you. Everyday AI helps you keep up and get ahead.Listen on: Apple Podcasts   SpotifySupport the showMore at https://linktr.ee/EvanKirstel

Episode metadata supplied by the publisher feed · Published Dec 15, 2025

Embed this episode

Interested in being a guest? Email us at [email protected] Big leaps in AI rarely come from one breakthrough. They emerge when hardware design, open software, and real workloads click into place. That’s the story we unpack with AMD’s Ramine Roane: how an open, developer-first approach combined with high-bandwidth memory, chiplet packaging, and a re-architected software stack is reshaping performance and cost from the edge to the largest data centers. We walk through why memory capacity a...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

Inside AMD’s AI Strategy From Edge To Data Center

0:00 30:18

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of What's Up with Tech??

This episode is 30 minutes long.

When was this What's Up with Tech? episode published?

This episode was published on December 15, 2025.

Is there a transcript available for this episode?

Yes, a full transcript is available for this episode. You can read the complete transcript on the episode page.

Can I download this What's Up with Tech? episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!