TurboQuant: Google's 6x KV Cache Compression and the Quiet Economics of Long Context AI - June 14, 2026 episode artwork

EPISODE · Jun 14, 2026 · 12 MIN

TurboQuant: Google's 6x KV Cache Compression and the Quiet Economics of Long Context AI - June 14, 2026

from DX Today | No-Hype Podcast & News About AI & DX

TurboQuant: Google's 6x KV Cache Compression and the Quiet Economics of Long Context AI Google Research's TurboQuant compresses the LLM key value cache to roughly three bits per coordinate with near zero accuracy loss, delivering at least six times less memory and up to eight times faster attention on NVIDIA H100 GPUs. We unpack how its two stage design pairs a training free random rotation with a one bit correction step, why a 70B model's 128K context cache shrinks from about 40GB to under 7GB, and what that means for the cost of long context AI everywhere. Hosted by Rick Spair and Laura. The DX Today Podcast brings you daily deep dives into the most consequential stories in the AI ecosystem. Send us fan mail: https://dxtoday.com/contact #AI #LLMInference #KVCache #Quantization #TechNews

Episode metadata supplied by the publisher feed · Published Jun 14, 2026

Embed this episode

TurboQuant: Google's 6x KV Cache Compression and the Quiet Economics of Long Context AI Google Research's TurboQuant compresses the LLM key value cache to roughly three bits per coordinate with near zero accuracy loss, delivering at least six times less memory and up to eight times faster attention on NVIDIA H100 GPUs. We unpack how its two stage design pairs a training free random rotation with a one bit correction step, why a 70B model's 128K context cache shrinks from about 40GB to under ...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

TurboQuant: Google's 6x KV Cache Compression and the Quiet Economics of Long Context AI - June 14, 2026

0:00 12:47

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of DX Today | No-Hype Podcast & News About AI & DX?

This episode is 12 minutes long.

When was this DX Today | No-Hype Podcast & News About AI & DX episode published?

This episode was published on June 14, 2026.

Can I download this DX Today | No-Hype Podcast & News About AI & DX episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!