Serving MoE Models with Disaggregated Expert Parallelism episode artwork

EPISODE · May 19, 2026

Serving MoE Models with Disaggregated Expert Parallelism

from AI Post Transformers

This episode explores MegaScale-Infer, a systems paper on serving large mixture-of-experts language models by separating the attention path from the expert feed-forward path and scheduling them independently. It explains why MoE models can look efficient on paper yet still waste GPU capacity in practice, especially during decode, where KV-cache-heavy attention and uneven expert routing create very different bottlenecks. The discussion focuses on the paper’s core argument for disaggregated expert parallelism and a ping-pong microbatch pipeline designed to keep both attention and expert GPUs busy instead of leaving one side idle. Listeners would find it interesting for its clear look at the gap between model architecture and real-world serving performance, including a pointed debate over whether strong decode benchmarks actually translate into better end-to-end user latency. Sources: 1. MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism — Ruidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu, Cesar A. Stuardo, Dongyang Wang, Xinlei Zhang, Huaping Zhou, Haoran Wei, Yang Cheng, Jianzhe Xiao, Xinyi Zhang, Lingjun Liu, Haibin Lin, Li-Wen Chang, Jianxi Ye, Xiao Yu, Xuanzhe Liu, Xin Jin, Xin Liu, 2025 http://arxiv.org/abs/2504.02263 2. GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism — Yanping Huang, Yonglong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Miaosen Wang, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Zhifeng Chen, 2019 https://scholar.google.com/scholar?q=GPipe%3A+Efficient+Training+of+Giant+Neural+Networks+using+Pipeline+Parallelism 3. PipeDream: Generalized Pipeline Parallelism for DNN Training — Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil Devanur, Greg Ganger, Phil Gibbons, Matei Zaharia, 2019 https://scholar.google.com/scholar?q=PipeDream%3A+Generalized+Pipeline+Parallelism+for+DNN+Training 4. Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM — Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Nitin Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Matei Zaharia, 2021 https://scholar.google.com/scholar?q=Efficient+Large-Scale+Language+Model+Training+on+GPU+Clusters+Using+Megatron-LM 5. Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache — Bin Lin, Chen Zhang, Tao Peng, Hanyu Zhao, Wencong Xiao, Minmin Sun, Anmin Liu, Zhipeng Zhang, Lanbo Li, Xiafei Qiu, Shen Li, Zhigang Ji, Tao Xie, Yong Li, Wei Lin, 2024 https://scholar.google.com/scholar?q=Infinite-LLM%3A+Efficient+LLM+Service+for+Long+Context+with+DistAttention+and+Distributed+KVCache 6. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024 https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving 7. Splitwise: Efficient Generative LLM Inference using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Inigo Goiri, Saeed Maleki, Ricardo Bianchini, 2023 https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+using+Phase+Splitting 8. MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs — Shiyi Cao, Shu Liu, Tyler Griggs, Peter Schafhalter, Xiaoxuan Liu, Ying Sheng, Joseph E. Gonzalez, Matei Zaharia, Ion Stoica, 2024 https://scholar.google.com/scholar?q=MoE-Lightning%3A+High-Throughput+MoE+Inference+on+Memory-constrained+GPUs 9. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI et al., 2024 https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model 10. Toward Efficient Inference for Mixture of Experts — Haiyang Huang, Newsha Ardalani, Anna Sun, Liu Ke, Hsien-Hsin S. Lee, Shruti Bhosale, Carole-Jean Wu, Benjamin Lee, 2024 https://scholar.google.com/scholar?q=Toward+Efficient+Inference+for+Mixture+of+Experts 11. AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference — Shuzhang Zhong, Ling Liang, Yuan Wang, Runsheng Wang, Ru Huang, Meng Li, 2024 https://scholar.google.com/scholar?q=AdapMoE%3A+Adaptive+Sensitivity-based+Expert+Gating+and+Management+for+Efficient+MoE+Inference 12. HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference — Shuzhang Zhong, Yanfan Sun, Ling Liang, Runsheng Wang, Ru Huang, Meng Li, 2025 https://scholar.google.com/scholar?q=HybriMoE%3A+Hybrid+CPU-GPU+Scheduling+and+Cache+Management+for+Efficient+MoE+Inference 13. Oracle-MoE: Locality-preserving Routing in the Oracle Space for Memory-constrained Large Language Model Inference — Jixian Zhou, Fang Dong, Ruijun Huang, Hengjie Cao, Mengyi Chen, Yifeng Yang, Anrui Chen, Mingzhi Dong, Yujiang Wang, Dongsheng Li, David A. Clifton, Qin Lv, Rui Zhu, Chun Zhang, Fan Yang, Tun Lu, Ning Gu, Li Shang, 2025 https://scholar.google.com/scholar?q=Oracle-MoE%3A+Locality-preserving+Routing+in+the+Oracle+Space+for+Memory-constrained+Large+Language+Model+Inference 14. Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism — Xinglin Pan, Shaohuai Shi, Wenxiang Lin, Yuxin Wang, Zhenheng Tang, Wei Wang, Xiaowen Chu, 2025 https://scholar.google.com/scholar?q=Efficient+MoE+Inference+with+Fine-Grained+Scheduling+of+Disaggregated+Expert+Parallelism 15. AI Post Transformers: JANUS for Scalable MoE Inference — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-15-janus-for-scalable-moe-inference-78ae30.mp3 16. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3 17. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3 18. AI Post Transformers: NanoFlow and the Future of LLM Serving — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-15-nanoflow-and-the-future-of-llm-serving-7429c9.mp3 19. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3 20. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3 21. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3 22. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3 Interactive Visualization: Serving MoE Models with Disaggregated Expert Parallelism

Episode metadata supplied by the publisher feed · Published May 19, 2026

Embed this episode

NOW PLAYING

Serving MoE Models with Disaggregated Expert Parallelism

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on May 19, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!