Offloading LLM Attention: Q-Shipping and KV-Side Compute episode artwork

EPISODE · Oct 16, 2025 · 42 MIN

Offloading LLM Attention: Q-Shipping and KV-Side Compute

from The Gist Talk · host kw

The source provides an extensive overview of strategies, collectively termed Q-shipping and KV-side compute, aimed at overcoming the memory bandwidth bottleneck during Large Language Model (LLM) inference, particularly in the decode phase

Episode metadata supplied by the publisher feed · Published Oct 16, 2025

Embed this episode

NOW PLAYING

Offloading LLM Attention: Q-Shipping and KV-Side Compute

0:00 42:25

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of The Gist Talk?

This episode is 42 minutes long.

When was this The Gist Talk episode published?

This episode was published on October 16, 2025.

Can I download this The Gist Talk episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!