ForkKV for Multi-LoRA Agent Serving episode artwork

EPISODE · May 12, 2026

ForkKV for Multi-LoRA Agent Serving

from AI Post Transformers

This episode explores ForkKV, a systems paper on serving multiple LoRA-based agents from one base language model without duplicating massive KV caches for shared context. It explains why ordinary prefix caching breaks once different LoRA adapters change the activations, then walks through the paper’s core idea: split cache state into a large shared base and a small adapter-specific residual, using an operating-system-style copy-on-write model for agent branches. The discussion connects that design to prior work on LoRA, prefix caching, PagedAttention, and disaggregated memory, making the argument that the real win is practical GPU memory efficiency for coding assistants and tool-using agent workflows. Listeners would find it interesting because it frames transformer serving as a memory-management problem and shows how borrowing ideas from Unix process forking could make multi-agent LLM systems far more scalable. Sources: 1. ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache — Shao Wang, Rui Ren, Lin Gui, 2026 http://arxiv.org/abs/2604.06370 2. The UNIX Time-Sharing System — Dennis M. Ritchie and Ken Thompson, 1974 https://scholar.google.com/scholar?q=The+UNIX+Time-Sharing+System 3. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica, 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 4. vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention — Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, and Ashish Panwar, 2024 https://scholar.google.com/scholar?q=vAttention%3A+Dynamic+Memory+Management+for+Serving+LLMs+without+PagedAttention 5. ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache — Shao Wang, Rui Ren, and Lin Gui, 2026 https://scholar.google.com/scholar?q=ForkKV%3A+Scaling+Multi-LoRA+Agent+Serving+via+Copy-on-Write+Disaggregated+KV+Cache 6. Efficiently Programming Large Language Models using SGLang — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng, 2023 https://scholar.google.com/scholar?q=Efficiently+Programming+Large+Language+Models+using+SGLang 7. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang, 2024 https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving 8. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu, Heyang Huang, Junhao Hu, Jiang Xu, Xusheng Chen, Tao Xie, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, and Yizhou Shan, 2024 https://scholar.google.com/scholar?q=MemServe%3A+Context+Caching+for+Disaggregated+LLM+Serving+with+Elastic+Memory+Pool 9. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu, 2024 https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving 10. LRAgent: efficient kv cache sharing for multi-lora llm agents — H. Jeon, H. Ha, and J. Kim, 2026 https://scholar.google.com/scholar?q=LRAgent%3A+efficient+kv+cache+sharing+for+multi-lora+llm+agents 11. Punica: multi-tenant lora serving — L. Chen, Z. Ye, Y. Wu, D. Zhuo, L. Ceze, and A. Krishnamurthy, 2024 https://scholar.google.com/scholar?q=Punica%3A+multi-tenant+lora+serving 12. S-LoRA: serving thousands of concurrent lora adapters — Y. Sheng, S. Cao, D. Li, C. Hooper, N. Lee, S. Yang, C. Chou, B. Zhu, L. Zheng, K. Keutzer, J. E. Gonzalez, and I. Stoica, 2023 https://scholar.google.com/scholar?q=S-LoRA%3A+serving+thousands+of+concurrent+lora+adapters 13. DLoRA: dynamically orchestrating requests and adapters for LoRA LLM serving — B. Wu, R. Zhu, Z. Zhang, P. Sun, X. Liu, and X. Jin, 2024 https://scholar.google.com/scholar?q=DLoRA%3A+dynamically+orchestrating+requests+and+adapters+for+LoRA+LLM+serving 14. Tokencake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications — Z. Bian, F. Wu, T. Ma, and Y. Zhuo, 2025 https://scholar.google.com/scholar?q=Tokencake%3A+A+KV-Cache-centric+Serving+Framework+for+LLM-based+Multi-Agent+Applications 15. KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows — Z. Pan, A. Patel, Z. Hu, Y. Shen, Y. Guan, W. Li, L. Qin, Y. Wang, and Y. Ding, 2025 https://scholar.google.com/scholar?q=KVFlow%3A+Efficient+Prefix+Caching+for+Accelerating+LLM-Based+Multi-Agent+Workflows 16. MobiLoRA: Accelerating LoRA-based LLM Inference on Mobile Devices via Context-aware KV Cache Optimization — Borui Li et al., 2025 https://scholar.google.com/scholar?q=MobiLoRA%3A+Accelerating+LoRA-based+LLM+Inference+on+Mobile+Devices+via+Context-aware+KV+Cache+Optimization 17. Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA — Allison Li, Kristjan Greenewald, Thomas Parnell, Navid Azizan, 2025 https://scholar.google.com/scholar?q=Efficient+Multi-Adapter+LLM+Serving+via+Cross-Model+KV-Cache+Reuse+with+Activated+LoRA 18. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025 https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse 19. AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization — Qiyang Li et al., 2026 https://scholar.google.com/scholar?q=AdaFuse%3A+Accelerating+Dynamic+Adapter+Inference+via+Token-Level+Pre-Gating+and+Fused+Kernel+Optimization 20. ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference — Xiang Liu et al., 2025 https://scholar.google.com/scholar?q=ChunkKV%3A+Semantic-Preserving+KV+Cache+Compression+for+Efficient+Long-Context+LLM+Inference 21. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yucheng Li et al., 2025 https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods 22. ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM Serving — Jiuchen Shi et al., 2026 https://scholar.google.com/scholar?q=ELORA%3A+Efficient+LoRA+and+KV+Cache+Management+for+Multi-LoRA+LLM+Serving 23. AI Post Transformers: Efficient KV Cache Sharing for Multi-LoRA Agents — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-22-efficient-kv-cache-sharing-for-multi-lor-afda05.mp3 24. AI Post Transformers: TokenDance for Multi-Agent KV Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-22-tokendance-for-multi-agent-kv-cache-shar-aa9b99.mp3 25. AI Post Transformers: ContiguousKV for Faster LLM Prefill KV Reuse — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-20-contiguouskv-for-faster-llm-prefill-kv-r-59f545.mp3 26. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3 27. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3 Interactive Visualization: ForkKV for Multi-LoRA Agent Serving

Episode metadata supplied by the publisher feed · Published May 12, 2026

Embed this episode

NOW PLAYING

ForkKV for Multi-LoRA Agent Serving

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on May 12, 2026.

Is there a transcript available for this episode?

Yes, a full transcript is available for this episode. You can read the complete transcript on the episode page.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!