MegaScale-Infer: Disaggregating Experts for Faster MoE Serving episode artwork

EPISODE · Jul 26, 2026

MegaScale-Infer: Disaggregating Experts for Faster MoE Serving

from AI Post Transformers

This episode explores MegaScale-Infer, a ByteDance Seed/Peking University system for serving large Mixture-of-Experts models more efficiently during inference. The discussion breaks down why decode-phase attention is memory-bandwidth-bound rather than compute-bound, and how MoE's top-k expert routing—while cutting theoretical FLOPs—actually shrinks the effective batch size each expert sees, tanking GPU utilization (illustrated with a concrete Mixtral 8x22B example dropping to 25% utilization). The core proposed fix is disaggregation: physically separating attention and expert computation onto independently scaled GPU pools so attention replicas can pool enough requests to keep experts saturated, building on prior work like DistServe's prefill/decode split. Listeners interested in the gap between algorithmic efficiency claims and real-world GPU serving costs will find the roofline-model analysis a sharp corrective to "sparsity equals free lunch" thinking. Sources: 1. MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism — Ruidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu, Cesar A. Stuardo, Dongyang Wang, Xinlei Zhang, Huaping Zhou, Haoran Wei, Yang Cheng, Jianzhe Xiao, Xinyi Zhang, Lingjun Liu, Haibin Lin, Li-Wen Chang, Jianxi Ye, Xiao Yu, Xuanzhe Liu, Xin Jin, Xin Liu, 2025 http://arxiv.org/abs/2504.02263v1 2. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024 https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+Large+Language+Model+Serving 3. Mooncake: Kimi's KVCache-centric Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2024 https://scholar.google.com/scholar?q=Mooncake%3A+Kimi%27s+KVCache-centric+Architecture+for+LLM+Serving 4. DeepSeek-V3 Technical Report — DeepSeek-AI (Aixin Liu et al.), 2024 https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report 5. MoE-Lightning: High-Throughput MoE Inference on Memory-Constrained GPUs — Shiyi Cao, Shu Liu, Tyler Griggs, Peter Schafhalter, Xiaoxuan Liu, Ying Sheng, Joseph E. Gonzalez, Matei Zaharia, Ion Stoica, 2024 https://scholar.google.com/scholar?q=MoE-Lightning%3A+High-Throughput+MoE+Inference+on+Memory-Constrained+GPUs 6. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding — Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, Zhifeng Chen, 2020 https://scholar.google.com/scholar?q=GShard%3A+Scaling+Giant+Models+with+Conditional+Computation+and+Automatic+Sharding Interactive Visualization: MegaScale-Infer: Disaggregating Experts for Faster MoE Serving

Episode metadata supplied by the publisher feed · Published Jul 26, 2026

Embed this episode

NOW PLAYING

MegaScale-Infer: Disaggregating Experts for Faster MoE Serving

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on July 26, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!