Mooncake for KV Cache-Centric LLM Serving episode artwork

EPISODE · Jun 5, 2026

Mooncake for KV Cache-Centric LLM Serving

from AI Post Transformers

This episode explores Mooncake, a production LLM serving architecture that treats KV cache reuse and movement as the central challenge in long-context chat, not just raw GPU compute. It explains why prefill and decode stress hardware in different ways, how metrics like time to first token and time between tokens drive system design, and why separating those phases helps meet real latency targets. The discussion walks through Mooncake’s cache-first scheduler, tiered KV storage across GPU memory, CPU DRAM, and SSD, and its use of RDMA, chunked prefill, and layer-wise overlap to start decoding sooner while reusing existing state. It also argues that Mooncake’s interest lies less in a single breakthrough than in how it combines prefix-aware routing, overload-aware early rejection, and cross-node KV reuse into a practical serving stack for large-scale chat systems. Sources: 1. Mooncake for KV Cache-Centric LLM Serving https://arxiv.org/pdf/2407.00079 2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, et al., 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 3. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Ying Sheng, et al., 2023 https://scholar.google.com/scholar?q=SGLang%3A+Efficient+Execution+of+Structured+Language+Model+Programs 4. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024 https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving 5. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2024 https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving 6. Overload Control for Scaling WeChat Microservices — Hao Zhou, Ming Chen, Qian Lin, Yong Wang, Xiaobin She, Sifan Liu, Rui Gu, Beng Chin Ooi, Junfeng Yang, 2018 https://scholar.google.com/scholar?q=Overload+Control+for+Scaling+WeChat+Microservices 7. Overload Control for microsecond-scale RPCs with Breakwater — Inho Cho, Ahmed Saeed, Joshua Fried, Seo Jin Park, Mohammad Alizadeh, Adam Belay, 2020 https://scholar.google.com/scholar?q=Overload+Control+for+microsecond-scale+RPCs+with+Breakwater 8. SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference — Yinghao Tang, Tingfeng Lan, Xiuqi Huang, Hui Lu, Wei Chen, 2025 https://scholar.google.com/scholar?q=SCORPIO%3A+Serving+the+Right+Requests+at+the+Right+Time+for+Heterogeneous+SLOs+in+LLM+Inference 9. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, Inigo Goiri, Aashaka Shah, Saeed Maleki, Ricardo Bianchini, 2023 https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting 10. AttentionStore: Cost-Effective Attention Reuse Across Multi-Turn Conversations in Large Language Model Serving — Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, Pengfei Zuo, 2024 https://scholar.google.com/scholar?q=AttentionStore%3A+Cost-Effective+Attention+Reuse+Across+Multi-Turn+Conversations+in+Large+Language+Model+Serving 11. Preble: Efficient Distributed Prompt Scheduling for LLM Serving — Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, Yiying Zhang, 2024 https://scholar.google.com/scholar?q=Preble%3A+Efficient+Distributed+Prompt+Scheduling+for+LLM+Serving 12. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, Junchen Jiang, 2024 https://scholar.google.com/scholar?q=CacheBlend%3A+Fast+Large+Language+Model+Serving+for+RAG+with+Cached+Knowledge+Fusion 13. P/D-Serve: Serving Disaggregated Large Language Model at Scale — Yibo Jin et al., 2024 https://scholar.google.com/scholar?q=P%2FD-Serve%3A+Serving+Disaggregated+Large+Language+Model+at+Scale 14. POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference — Aditya K. Kamath et al., 2025 https://scholar.google.com/scholar?q=POD-Attention%3A+Unlocking+Full+Prefill-Decode+Overlap+for+Faster+LLM+Inference 15. Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models — Siyan Zhao et al., 2024 https://scholar.google.com/scholar?q=Prepacking%3A+A+Simple+Method+for+Fast+Prefilling+and+Increased+Throughput+in+Large+Language+Models 16. Mustafar: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference — Donghyeon Joo et al., 2025 https://scholar.google.com/scholar?q=Mustafar%3A+Promoting+Unstructured+Sparsity+for+KV+Cache+Pruning+in+LLM+Inference 17. SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning — Huanxuan Liao et al., 2025/2026 https://scholar.google.com/scholar?q=SparK%3A+Query-Aware+Unstructured+Sparsity+with+Recoverable+KV+Cache+Channel+Pruning 18. More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression — Jiebin Zhang et al., 2024 https://scholar.google.com/scholar?q=More+Tokens%2C+Lower+Precision%3A+Towards+the+Optimal+Token-Precision+Trade-off+in+KV+Cache+Compression 19. KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse — Huan Yang et al., 2025 https://scholar.google.com/scholar?q=KVShare%3A+An+LLM+Service+System+with+Efficient+and+Effective+Multi-Tenant+KV+Cache+Reuse 20. CacheSolidarity: Preventing Prefix Caching Side Channels in Multi-tenant LLM Serving Systems — Panagiotis Georgios Pennas et al., 2026 https://scholar.google.com/scholar?q=CacheSolidarity%3A+Preventing+Prefix+Caching+Side+Channels+in+Multi-tenant+LLM+Serving+Systems 21. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3 22. AI Post Transformers: Characterizing LLM KV Cache Workloads in Production — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/characterizing-llm-kv-cache-workloads-in-production/ 23. AI Post Transformers: ContiguousKV for Faster LLM Prefill KV Reuse — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-20-contiguouskv-for-faster-llm-prefill-kv-r-59f545.mp3 24. AI Post Transformers: CacheFlow and 3D-Parallel KV Cache Restoration — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-01-cacheflow-and-3d-parallel-kv-cache-resto-8db883.mp3 25. AI Post Transformers: Beluga: CXL Memory Pooling for LLM KV Cache — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-27-beluga-cxl-memory-pooling-for-llm-kv-cac-b6142f.mp3 26. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3 Interactive Visualization: Mooncake for KV Cache-Centric LLM Serving

Episode metadata supplied by the publisher feed · Published Jun 5, 2026

Embed this episode

NOW PLAYING

Mooncake for KV Cache-Centric LLM Serving

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 5, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!