The Practitioner's Guide to TurboQuant episode artwork

EPISODE · Apr 3, 2026 · 25 MIN

The Practitioner's Guide to TurboQuant

from The Weight Update · host Kris Moore

KV cache compression on your own hardware: what works, what doesn't, and when to care.Google's TurboQuant paper compresses KV cache to 3 bits per coordinate — 6x memory reduction, 8x faster inference, zero accuracy loss, no retraining required. This deep-dive walks through what it actually is, the three-layer compression stack, real benchmark results on a consumer RTX 4090, community implementations available today, the Hugging Face ecosystem integration, and a CTO decision framework for when this matters to your org. Companion to the LinkedIn article of the same name.25 sources cited. Full source list in show notes.AI Disclosure: This episode was produced with AI assistance. Research synthesis and script writing used Claude (Anthropic) under human editorial direction. Audio narration by Microsoft Edge TTS (en-US-AndrewNeural voice).

Episode metadata supplied by the publisher feed · Published Apr 3, 2026

Embed this episode

Ready to play

The Practitioner's Guide to TurboQuant

0:00 25:14

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of The Weight Update?

This episode is 25 minutes long.

When was this The Weight Update episode published?

This episode was published on April 3, 2026.

Can I download this The Weight Update episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!