PackKV Lossy Compression for KV Caches episode artwork

EPISODE · May 4, 2026

PackKV Lossy Compression for KV Caches

from AI Post Transformers

This episode explores PackKV, a method for shrinking the transformer KV cache during long-context inference by combining low-bit quantization with GPU-friendly repacking and lossy compression. It explains why KV cache growth can dominate memory use in large models, using examples where cache size exceeds model weights, and frames the problem as a systems bottleneck driven more by memory traffic than raw computation. The discussion compares PackKV to prior approaches such as KV quantization, token pruning, and offloading to CPU memory, highlighting the paper’s argument that compression is only useful if decompression is tightly integrated into the inference pipeline. A listener would find it interesting because it turns a seemingly low-level optimization into a broader claim about how future long-context LLM performance may depend as much on memory layout and kernel design as on model architecture. Sources: 1. PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression — Bo Jiang, Taolue Yang, Youyuan Liu, Xubin He, Sheng Di, Sian Jin, 2025 http://arxiv.org/abs/2512.24449 2. Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time — Zichang Liu, Aditya Desai, Fangshuo Liao, Victor Xie, Zhaozhuo Xu, Anastasios Kyrillidis, Anshumali Shrivastava, et al., 2023 https://scholar.google.com/scholar?q=Scissorhands%3A+Exploiting+the+Persistence+of+Importance+Hypothesis+for+LLM+KV+Cache+Compression+at+Test+Time 3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Zhao Song, Yuandong Tian, Clark Barrett, Zhangyang Wang, Beidi Chen, et al., 2023 https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models 4. KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache — Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, Xia Hu, 2024 https://scholar.google.com/scholar?q=KIVI%3A+A+Tuning-Free+Asymmetric+2bit+Quantization+for+KV+Cache 5. SnapKV: LLM Knows What You are Looking for Before Generation — Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, Deming Chen, 2024 https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+are+Looking+for+Before+Generation 6. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuyang Liu, Haotian Li, Yao Cheng, Siddhant Ray, Yizhou Huang, Qizhen Zhang, Kaixiang Du, Jinyang Yao, Shan Lu, Ganesh Ananthanarayanan et al., 2024 https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving 7. Q-Hitter: A Better Token Oracle for Efficient LLM Inference via Sparse-Quantized KV Cache — Zhuodong Zhang, Shang Liu, Ruobing Chen, Bhavya Kailkhura, Ben Chen, An Wang, 2024 https://scholar.google.com/scholar?q=Q-Hitter%3A+A+Better+Token+Oracle+for+Efficient+LLM+Inference+via+Sparse-Quantized+KV+Cache 8. PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling — Zefan Cai, Yichi Zhang, Bofei Gao, Yuliang Liu, Tianyu Liu, Keming Lu, Wayne Xiong, Yue Dong, Baobao Chang, Junjie Hu, Wen Xiao, 2025 https://scholar.google.com/scholar?q=PyramidKV%3A+Dynamic+KV+Cache+Compression+based+on+Pyramidal+Information+Funneling 9. Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution — Alessio Devoto, Maximilian Jeblick, Simon Jegou, 2025 https://scholar.google.com/scholar?q=Expected+Attention%3A+KV+Cache+Compression+by+Estimating+Attention+from+Future+Queries+Distribution 10. TurboQuant: Online Vector Quantization with Near-Optimal Distortion — Amir Zandieh, Majid Daliri, Majid Hadian, Vahab Mirrokni, 2025 https://scholar.google.com/scholar?q=TurboQuant%3A+Online+Vector+Quantization+with+Near-Optimal+Distortion 11. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025 https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse 12. HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse — Yuwei An et al., 2025 https://scholar.google.com/scholar?q=HyperRAG%3A+Enhancing+Quality-Efficiency+Tradeoffs+in+Retrieval-Augmented+Generation+with+Reranker+KV-Cache+Reuse 13. ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation — Shihao Wang et al., 2026 https://scholar.google.com/scholar?q=ProphetKV%3A+User-Query-Driven+Selective+Recomputation+for+Efficient+KV+Cache+Reuse+in+Retrieval-Augmented+Generation 14. ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification — Yefei He et al., 2024 https://scholar.google.com/scholar?q=ZipCache%3A+Accurate+and+Efficient+KV+Cache+Quantization+with+Salient+Token+Identification 15. KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs — Zunhai Su and Kehong Yuan, 2025 https://scholar.google.com/scholar?q=KVSink%3A+Understanding+and+Enhancing+the+Preservation+of+Attention+Sinks+in+KV+Cache+Quantization+for+LLMs 16. ThinK: Thinner Key Cache by Query-Driven Pruning — Yuhui Xu et al., 2024 https://scholar.google.com/scholar?q=ThinK%3A+Thinner+Key+Cache+by+Query-Driven+Pruning 17. KV-Compress: Paged KV-Cache Compression with Variable Compression Rates per Attention Head — Isaac Rehg, 2024 https://scholar.google.com/scholar?q=KV-Compress%3A+Paged+KV-Cache+Compression+with+Variable+Compression+Rates+per+Attention+Head 18. Paged Attention Meets FlexAttention: Unlocking Long-Context Efficiency in Deployed Inference — Thomas Joshi et al., 2025 https://scholar.google.com/scholar?q=Paged+Attention+Meets+FlexAttention%3A+Unlocking+Long-Context+Efficiency+in+Deployed+Inference 19. AI Post Transformers: TokenDance for Multi-Agent KV Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-22-tokendance-for-multi-agent-kv-cache-shar-aa9b99.mp3 20. AI Post Transformers: KVSwap for Disk-Aware Long-Context On-Device Inference — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-16-kvswap-for-disk-aware-long-context-on-de-f3c15e.mp3 21. AI Post Transformers: CacheFlow and 3D-Parallel KV Cache Restoration — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-01-cacheflow-and-3d-parallel-kv-cache-resto-8db883.mp3 22. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3 23. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3 24. AI Post Transformers: Computation-Bandwidth-Memory Trade-offs for AI Infrastructure — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-09-computation-bandwidth-memory-trade-offs-a83f2b.mp3 Interactive Visualization: PackKV Lossy Compression for KV Caches

Episode metadata supplied by the publisher feed · Published May 4, 2026

Embed this episode

NOW PLAYING

PackKV Lossy Compression for KV Caches

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on May 4, 2026.

Is there a transcript available for this episode?

Yes, a full transcript is available for this episode. You can read the complete transcript on the episode page.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!