EPISODE · May 2, 2026
DeltaKV: Compressing KV Caches for Long Context
from AI Post Transformers
This episode explores DeltaKV, a method for reducing the huge GPU memory burden of KV caches in long-context language model inference without simply discarding old tokens. It contrasts three strategies for handling long contexts: token eviction, dynamic sparse attention, and true compression, arguing that the cache contains structured redundancy that can be exploited rather than treated as disposable overhead. The discussion highlights DeltaKV’s core idea of keeping a small uncompressed reference set and storing other cache entries as compressed residuals relative to similar past states, drawing an analogy to delta encoding or version control. Listeners would find it interesting because it connects transformer internals, systems constraints, and practical serving performance, including claims of cutting memory to 29 percent of baseline and reaching up to 2x throughput with supporting infrastructure like Sparse-vLLM. Sources: 1. DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity — Jitai Hao, Qiang Huang, Yaowei Wang, Min Zhang, Jun Yu, 2026 http://arxiv.org/abs/2602.08005 2. Generalization through Memorization: Nearest Neighbor Language Models — Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, Mike Lewis, 2019 https://scholar.google.com/scholar?q=Generalization+through+Memorization%3A+Nearest+Neighbor+Language+Models 3. Reformer: The Efficient Transformer — Nikita Kitaev, Lukasz Kaiser, Anselm Levskaya, 2020 https://scholar.google.com/scholar?q=Reformer%3A+The+Efficient+Transformer 4. Improving Language Models by Retrieving from Trillions of Tokens — Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann and many others, 2021 https://scholar.google.com/scholar?q=Improving+Language+Models+by+Retrieving+from+Trillions+of+Tokens 5. Memorizing Transformers — Yuhuai Wu, Markus N. Rabe, DeLesley Hutchins, Christian Szegedy, 2022 https://scholar.google.com/scholar?q=Memorizing+Transformers 6. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu, Hanchen Li, Yihua Cheng and others, 2023 https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving 7. Palu: Compressing KV-Cache with Low-Rank Projection — Chi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin and others, 2024 https://scholar.google.com/scholar?q=Palu%3A+Compressing+KV-Cache+with+Low-Rank+Projection 8. Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries — Junhyuck Kim, Jongho Park, Jaewoong Cho, Dimitris Papailiopoulos, 2025 https://scholar.google.com/scholar?q=Lexico%3A+Extreme+KV+Cache+Compression+via+Sparse+Coding+over+Universal+Dictionaries 9. Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models — Alina Shutova, Vladimir Malinovskii, Vage Egiazarian and others, 2025 https://scholar.google.com/scholar?q=Cache+Me+If+You+Must%3A+Adaptive+Key-Value+Quantization+for+Large+Language+Models 10. OmniKV: Dynamic Context Selection for Efficient Long-Context LLMs — Jitai Hao, Yuke Zhu, Tian Wang, Jun Yu, Xin Xin, Bo Zheng, Zhaochun Ren, Sheng Guo, 2025 https://scholar.google.com/scholar?q=OmniKV%3A+Dynamic+Context+Selection+for+Efficient+Long-Context+LLMs 11. QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference — Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, Song Han, 2024 https://scholar.google.com/scholar?q=QUEST%3A+Query-Aware+Sparsity+for+Efficient+Long-Context+LLM+Inference 12. Palu: KV-Cache Compression with Low-Rank Projection — Chi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin, Chong-Yan Chen, Yu-Fang Hu, Pei-Shuo Wang, Ning-Chi Huang, Luis Ceze, Mohamed S. Abdelfattah, Kai-Chiang Wu, 2025 https://scholar.google.com/scholar?q=Palu%3A+KV-Cache+Compression+with+Low-Rank+Projection 13. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yucheng Li, Huiqiang Jiang, Qianhui Wu, Xufang Luo, Surin Ahn, Chengruidong Zhang, Amir H. Abdi, Dongsheng Li, Jianfeng Gao, Yuqing Yang, Lili Qiu, 2025 https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods 14. The Pitfalls of KV Cache Compression — Alex Chen, Renato Geh, Aditya Grover, Guy Van den Broeck, Daniel Israel, 2025 https://scholar.google.com/scholar?q=The+Pitfalls+of+KV+Cache+Compression 15. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, Hyun Oh Song, 2025 https://scholar.google.com/scholar?q=KVzip%3A+Query-Agnostic+KV+Cache+Compression+with+Context+Reconstruction 16. Kvlink: Accelerating Large Language Models via Efficient KV Cache Reuse — approx. recent LLM systems authors, 2025/2026 https://scholar.google.com/scholar?q=Kvlink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse 17. HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse — approx. recent RAG systems authors, 2025/2026 https://scholar.google.com/scholar?q=HyperRAG%3A+Enhancing+Quality-Efficiency+Tradeoffs+in+Retrieval-Augmented+Generation+with+Reranker+KV-Cache+Reuse 18. ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation — approx. recent RAG/serving authors, 2025/2026 https://scholar.google.com/scholar?q=ProphetKV%3A+User-Query-Driven+Selective+Recomputation+for+Efficient+KV+Cache+Reuse+in+Retrieval-Augmented+Generation 19. KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs — approx. recent KV quantization authors, 2025/2026 https://scholar.google.com/scholar?q=KVSink%3A+Understanding+and+Enhancing+the+Preservation+of+Attention+Sinks+in+KV+Cache+Quantization+for+LLMs 20. LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference — approx. recent long-context inference authors, 2025/2026 https://scholar.google.com/scholar?q=LLMs+Know+What+to+Drop%3A+Self-Attention+Guided+KV+Cache+Eviction+for+Efficient+Long-Context+Inference 21. Task-KV: Task-Aware KV Cache Optimization via Semantic Differentiation of Attention Heads — approx. recent attention/KV optimization authors, 2025/2026 https://scholar.google.com/scholar?q=Task-KV%3A+Task-Aware+KV+Cache+Optimization+via+Semantic+Differentiation+of+Attention+Heads 22. WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More — approx. recent quantization authors, 2025/2026 https://scholar.google.com/scholar?q=WKVQuant%3A+Quantizing+Weight+and+Key%2FValue+Cache+for+Large+Language+Models+Gains+More 23. AlignedKV: Reducing Memory Access of KV-Cache with Precision-Aligned Quantization — approx. recent KV systems/quantization authors, 2025/2026 https://scholar.google.com/scholar?q=AlignedKV%3A+Reducing+Memory+Access+of+KV-Cache+with+Precision-Aligned+Quantization 24. Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off — approx. recent sparse attention authors, 2025/2026 https://scholar.google.com/scholar?q=Making+Every+Head+Count%3A+Sparse+Attention+Without+the+Speed-Performance+Trade-off 25. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention — approx. recent sparse systems authors, 2025/2026 https://scholar.google.com/scholar?q=Native+Sparse+Attention%3A+Hardware-Aligned+and+Natively+Trainable+Sparse+Attention 26. AI Post Transformers: Stochastic KV Routing for Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-29-stochastic-kv-routing-for-cache-sharing-5fef63.mp3 27. AI Post Transformers: ContiguousKV for Faster LLM Prefill KV Reuse — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-20-contiguouskv-for-faster-llm-prefill-kv-r-59f545.mp3 28. AI Post Transformers: KVSwap for Disk-Aware Long-Context On-Device Inference — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-16-kvswap-for-disk-aware-long-context-on-de-f3c15e.mp3 29. AI Post Transformers: Native Sparse Attention: Efficient Long-Context LLMs — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/native-sparse-attention-efficient-long-context-llms/ 30. AI Post Transformers: Memory Sparse Attention for 100M-Token Scaling — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-07-memory-sparse-attention-for-100m-token-s-377cff.mp3 31. AI Post Transformers: AWQ: On-Device LLM Compression and Acceleration — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/awq-on-device-llm-compression-and-acceleration/
Embed this episode
NOW PLAYING
DeltaKV: Compressing KV Caches for Long Context
No transcript for this episode yet
Similar Episodes
No similar episodes found.