VeriCache: Lossless LLM Inference from Lossy KV Caches episode artwork

EPISODE · Jun 5, 2026

VeriCache: Lossless LLM Inference from Lossy KV Caches

from AI Post Transformers

This episode explores VeriCache, a systems paper that asks whether a large language model can draft tokens using a compressed, lossy KV cache and then verify them against the full cache to recover exactly the same greedy-decoding output. It explains why KV cache size has become a central bottleneck in long-context inference, unpacking token dropping, KV quantization, prefix caching, and the speculative decoding ideas that VeriCache turns into a closed-loop draft-and-verify scheme. The discussion argues that fluency is not enough for real deployments because even a single wrong token can break code, structured JSON, or tool calls, so exact token-by-token agreement is the real standard for safe acceleration. Listeners get a clear picture of where serving stacks such as vLLM, Hugging Face TGI, and NVIDIA already use neighboring optimizations, and why VeriCache’s specific lossless-verification approach is both technically appealing and operationally difficult. Sources: 1. VeriCache: Turning Lossy KV Cache into Lossless LLM Inference — Jiayi Yao, Samuel Shen, Kuntai Du, Shaoting Feng, Dongjoo Seo, Rui Zhang, Yuyang Huang, Yuhan Liu, Shan Lu, Junchen Jiang, 2026 http://arxiv.org/abs/2605.17613 2. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023 https://proceedings.mlr.press/v202/leviathan23a.html 3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Beidi Chen, et al., 2023 https://proceedings.neurips.cc/paper_files/paper/2023/hash/6ceefa7b15572587b78ecfcebb2827f8-Abstract-Conference.html 4. KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache — Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Beidi Chen, Xia Hu, et al., 2024 https://arxiv.org/abs/2402.02750 5. QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache — Rishabh Tiwari, Haocheng Xi, Aditya Tomar, Coleman Hooper, Kurt Keutzer, Amir Gholami, et al., 2025 https://arxiv.org/abs/2502.10424 6. MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding — Ranajoy Sadhukhan, Jian Chen, Zhuoming Chen, et al., 2024 https://scholar.google.com/scholar?q=MagicDec%3A+Breaking+the+Latency-Throughput+Tradeoff+for+Long+Context+Generation+with+Speculative+Decoding 7. Accelerating Large-Scale Reasoning Model Inference: Self-Speculative Decoding with Sparse Attention — Yilong Zhao, Jiaming Tang, Kan Zhu, Zihao Ye, Chi-Chih Chang, Chaofan Lin, Jongseok Park, Guangxuan Xiao, Mohamed S. Abdelfattah, Mingyu Gao, Baris Kasikci, Song Han, and Ion Stoica, 2025 https://scholar.google.com/scholar?q=Accelerating+Large-Scale+Reasoning+Model+Inference%3A+Self-Speculative+Decoding+with+Sparse+Attention 8. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray, Yuyang Huang, Qizheng Zhang, Kuntai Du, Jiayi Yao, Shan Lu, Ganesh Ananthanarayanan, and Junchen Jiang, 2024 https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving 9. ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching — Xingyu Xiang, Raj Joshi, Yuhan Liu, Jiayi Yao, Chenxingyu Zhao, Junchen Jiang, Yang Zhou, Eddie Kohler, and Minlan Yu, 2025 https://scholar.google.com/scholar?q=ShadowServe%3A+Interference-Free+KV+Cache+Fetching+for+Distributed+Prefix+Caching 10. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yuhan Liu, Yihua Cheng, Jiayi Yao, Yuwei An, Xiaokun Chen, Shaoting Feng, et al., 2025 https://scholar.google.com/scholar?q=LMCache%3A+An+Efficient+KV+Cache+Layer+for+Enterprise-Scale+LLM+Inference 11. ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs — Jianlong Lei and Shashikant Ilager, 2026 https://arxiv.org/abs/2603.08727 12. Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs — Sayed Pedram Haeri Boroujeni et al., 2026 https://arxiv.org/abs/2604.04722 13. LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification — Penghui Yang et al., 2025 https://arxiv.org/abs/2502.17421 14. RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding — Guanzheng Chen et al., 2025 https://arxiv.org/abs/2502.20330 15. LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management — Yi Xiong et al., 2024 https://arxiv.org/abs/2410.00428 16. TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving — Bingyang Wu et al., 2025 https://arxiv.org/abs/2508.17219 17. KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows — Zaifeng Pan et al., 2025 https://arxiv.org/abs/2507.07400 18. KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse — Huan Yang et al., 2025 https://arxiv.org/abs/2503.16525 19. KV Cache Offloading for Context-Intensive Tasks — Andrey Bocharnikov et al., 2026 https://arxiv.org/abs/2604.08426 20. AI Post Transformers: KVzip for Query-Agnostic KV Cache Compression — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-29-kvzip-for-query-agnostic-kv-cache-compre-72afe5.mp3 21. AI Post Transformers: PackKV Lossy Compression for KV Caches — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-04-packkv-lossy-compression-for-kv-caches-b37bce.mp3 22. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3 23. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3 24. AI Post Transformers: Stochastic KV Routing for Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-29-stochastic-kv-routing-for-cache-sharing-5fef63.mp3

Episode metadata supplied by the publisher feed · Published Jun 5, 2026

Embed this episode

NOW PLAYING

VeriCache: Lossless LLM Inference from Lossy KV Caches

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 5, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!