EPISODE · Oct 16, 2025 · 42 MIN
Offloading LLM Attention: Q-Shipping and KV-Side Compute
from The Gist Talk · host kw
The source provides an extensive overview of strategies, collectively termed Q-shipping and KV-side compute, aimed at overcoming the memory bandwidth bottleneck during Large Language Model (LLM) inference, particularly in the decode phase
Embed this episode
NOW PLAYING
Offloading LLM Attention: Q-Shipping and KV-Side Compute
0:00
42:25
1×
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.
Frequently Asked Questions
How long is this episode of The Gist Talk?
This episode is 42 minutes long.
When was this The Gist Talk episode published?
This episode was published on October 16, 2025.
Can I download this The Gist Talk episode?
Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!