SE Radio 703: Sahaj Garg on Low Latency AI episode artwork

EPISODE · Jan 14, 2026 · 54 MIN

SE Radio 703: Sahaj Garg on Low Latency AI

from Software Engineering Radio - the podcast for professional software developers · host SE Radio

In this episode, Sahaj Garg, CTO of wispr.ai, joins SE Radio host Robert Blumen to talk about the challenges of building low-latency AI applications. They discuss latency's effect on consumer behavior as well as interactive applications. The conversation explores how to measure latency and how scale impacts it. Then Sahaj and Robert shift to themes around AI, including whether "AI" means LLMs or something broader, as they look at latency requirements and challenges around subtypes of AI applications. The final part of the episode explores techniques for managing latency in AI: speed vs accuracy trade-offs; speed vs cost; latency vs cost; choosing the right model; reducing quantization; distillation; and guessing + validating. Brought to you by IEEE Computer Society and IEEE Software magazine.

Episode metadata supplied by the publisher feed · Published Jan 14, 2026

Embed this episode

NOW PLAYING

SE Radio 703: Sahaj Garg on Low Latency AI

0:00 54:50

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Software Engineering Radio - the podcast for professional software developers?

This episode is 54 minutes long.

When was this Software Engineering Radio - the podcast for professional software developers episode published?

This episode was published on January 14, 2026.

Can I download this Software Engineering Radio - the podcast for professional software developers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!