Training Modular KV Caches at Scale episode artwork

EPISODE · Jun 17, 2026

Training Modular KV Caches at Scale

from AI Post Transformers

This episode explores the paper Cartridges at Scale, which asks whether large document collections can be distilled into reusable modular KV-cache memories so a model can answer questions without repeatedly rereading raw text. It explains what a cartridge is, how context distillation turns full-document context into compact learned prefixes, and why that differs from prompt caching, fine-tuning, ordinary long-context prompting, and text RAG. The discussion centers on the paper’s main claim that per-document memories do not reliably compose when trained independently, so the authors jointly train cartridges with both relevant and irrelevant memories present to teach a frozen model which compressed document to attend to in a noisy multi-document setting. Listeners would find it interesting because it treats the KV cache as a potential external memory layer that could reduce inference cost and latency while exposing hard questions about compositionality, transparency, and whether learned memory modules can outperform standard retrieval pipelines. Sources: 1. Cartridges at Scale: Training Modular KV Caches over Large Document Collections — Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert, 2026 http://arxiv.org/abs/2606.04557 2. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Douwe Kiela, et al., 2020 https://arxiv.org/abs/2005.11401 3. Prompt Cache: Modular Attention Reuse for Low-Latency Inference — In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, Lin Zhong, 2023 https://arxiv.org/abs/2311.04934 4. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study — Sabri Eyuboglu, Ryan Ehrlich, Simran Arora, Neel Guha, James Zou, Azalia Mirhoseini, Christopher Re, et al., 2025 https://arxiv.org/abs/2506.06266 5. Cartridges at Scale: Training Modular KV Caches over Large Document Collections — Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert, 2026 https://arxiv.org/abs/2606.04557 6. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study (https://arxiv.org/abs/2506.06266) — Sabri Eyuboglu, Ryan S. Ehrlich, Simran Arora, Neel Guha, Dylan Zinsley, Emily R. Liu, Atri Rudra, James Y. Zou, Azalia Mirhoseini, Christopher Re, 2025 https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+General-Purpose+Long+Context+Representations+via+Self-Study+%28https%3A%2F%2Farxiv.org%2Fabs%2F2506.06266%29 7. Learned Structure in CARTRIDGES: Keys as Shareable Routers in Self-Studied Representations (https://arxiv.org/abs/2508.17032) — Maurizio Diaz, 2025 https://scholar.google.com/scholar?q=Learned+Structure+in+CARTRIDGES%3A+Keys+as+Shareable+Routers+in+Self-Studied+Representations+%28https%3A%2F%2Farxiv.org%2Fabs%2F2508.17032%29 8. xRAG: Extreme Context Compression for Retrieval-Augmented Generation with One Token (https://arxiv.org/abs/2405.13792) — Xin Cheng, Xun Wang, Xingxing Zhang, Tao Ge, Si-Qing Chen, Furu Wei, Huishuai Zhang, Dongyan Zhao, 2024 https://scholar.google.com/scholar?q=xRAG%3A+Extreme+Context+Compression+for+Retrieval-Augmented+Generation+with+One+Token+%28https%3A%2F%2Farxiv.org%2Fabs%2F2405.13792%29 9. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction (https://arxiv.org/abs/2505.23416) — Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, Hyun Oh Song, 2025 https://scholar.google.com/scholar?q=KVzip%3A+Query-Agnostic+KV+Cache+Compression+with+Context+Reconstruction+%28https%3A%2F%2Farxiv.org%2Fabs%2F2505.23416%29 10. T2-RAGBench: Text-and-Table Benchmark for Evaluating Retrieval-Augmented Generation (https://aclanthology.org/2026.eacl-long.8/) — Jan Strich, Enes Kutay Isgorur, Maximilian Trescher, Chris Biemann, Martin Semmann, 2026 https://scholar.google.com/scholar?q=T2-RAGBench%3A+Text-and-Table+Benchmark+for+Evaluating+Retrieval-Augmented+Generation+%28https%3A%2F%2Faclanthology.org%2F2026.eacl-long.8%2F%29 11. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025 https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse 12. Hierarchical Document Refinement for Long-context Retrieval-augmented Generation — Jiajie Jin et al., 2025 https://scholar.google.com/scholar?q=Hierarchical+Document+Refinement+for+Long-context+Retrieval-augmented+Generation 13. LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs -- No Silver Bullet for LC or RAG Routing — Kuan Li et al., 2025 https://scholar.google.com/scholar?q=LaRA%3A+Benchmarking+Retrieval-Augmented+Generation+and+Long-Context+LLMs+--+No+Silver+Bullet+for+LC+or+RAG+Routing 14. ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities — Peng Xu et al., 2024 https://scholar.google.com/scholar?q=ChatQA+2%3A+Bridging+the+Gap+to+Proprietary+LLMs+in+Long+Context+and+RAG+Capabilities 15. LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain — Nicholas Pipitone and Ghita Houir Alami, 2024 https://scholar.google.com/scholar?q=LegalBench-RAG%3A+A+Benchmark+for+Retrieval-Augmented+Generation+in+the+Legal+Domain 16. KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse — Huan Yang et al., 2025 https://scholar.google.com/scholar?q=KVShare%3A+An+LLM+Service+System+with+Efficient+and+Effective+Multi-Tenant+KV+Cache+Reuse 17. AI Post Transformers: KVzip for Query-Agnostic KV Cache Compression — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-29-kvzip-for-query-agnostic-kv-cache-compre-72afe5.mp3 18. AI Post Transformers: From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-22-from-prefix-cache-to-fusion-rag-9c5d39.mp3 19. AI Post Transformers: Experimental Comparison of Agentic and Enhanced RAG — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-14-experimental-comparison-of-agentic-and-e-37d8bc.mp3 20. AI Post Transformers: Can Models Learn from Long Context? — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-can-models-learn-from-long-context-77533e.mp3 21. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp3 22. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3

Episode metadata supplied by the publisher feed · Published Jun 17, 2026

Embed this episode

NOW PLAYING

Training Modular KV Caches at Scale

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 17, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!