Moebius: Seamless Parallelism Switching for MoE Serving episode artwork

EPISODE · Jun 29, 2026

Moebius: Seamless Parallelism Switching for MoE Serving

from AI Post Transformers

This episode explores Moebius, a serving system for mixture-of-experts transformers that can switch at runtime between tensor parallelism and expert parallelism without restarting or draining live requests. It explains why tensor parallelism tends to give lower latency at low concurrency, while expert parallelism delivers better throughput at high concurrency, making bursty online traffic and RL rollouts natural settings where the best strategy changes over time. The discussion focuses on the hard systems problems behind that switch, including migrating in-flight requests, preserving paged KV caches, coping with CUDA graph address constraints, and handling KV-head mismatches that can waste cache capacity under tensor parallelism. It argues that the paper’s key contribution is treating the switch as a change in ownership and memory layout over one resident model and KV state, offering a concrete blueprint for serving large sparse models more efficiently. Sources: 1. Moebius: Serving Mixture-of-Expert Models with Seamless Runtime Parallelism Switch — Shaoyu Wang, Yizhuo Liang, Jaeyong Song, Chong Li, Seo Jin Park, 2026 http://arxiv.org/abs/2606.26607 2. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding — Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, et al., 2020 https://scholar.google.com/scholar?q=GShard%3A+Scaling+Giant+Models+with+Conditional+Computation+and+Automatic+Sharding 3. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity — William Fedus, Barret Zoph, Noam Shazeer, 2021 https://scholar.google.com/scholar?q=Switch+Transformers%3A+Scaling+to+Trillion+Parameter+Models+with+Simple+and+Efficient+Sparsity 4. DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale — Samyam Rajbhandari, Conglong Li, Zhewei Yao, et al., 2022 https://scholar.google.com/scholar?q=DeepSpeed-MoE%3A+Advancing+Mixture-of-Experts+Inference+and+Training+to+Power+Next-Generation+AI+Scale 5. MegaBlocks: Efficient Sparse Training with Mixture-of-Experts — Trevor Gale, Deepak Narayanan, Cliff Young, Matei Zaharia, 2022 https://scholar.google.com/scholar?q=MegaBlocks%3A+Efficient+Sparse+Training+with+Mixture-of-Experts 6. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism — Mohammad Shoeybi, Mostofa Patwary, Raul Puri, et al., 2019 https://scholar.google.com/scholar?q=Megatron-LM%3A+Training+Multi-Billion+Parameter+Language+Models+Using+Model+Parallelism 7. Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM — Deepak Narayanan, Mohammad Shoeybi, Jared Casper, et al., 2021 https://scholar.google.com/scholar?q=Efficient+Large-Scale+Language+Model+Training+on+GPU+Clusters+Using+Megatron-LM 8. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al., 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 9. Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism — Vikranth Srivatsa, Zijian He, Pu Guo, et al., 2026 https://scholar.google.com/scholar?q=Nitsum%3A+Serving+Tiered+LLM+Requests+with+Adaptive+Tensor+Parallelism 10. HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference — Haoran Lin et al., 2025 https://scholar.google.com/scholar?q=HAP%3A+Hybrid+Adaptive+Parallelism+for+Efficient+Mixture-of-Experts+Inference 11. Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services — Haoyu Chen et al., 2026 https://scholar.google.com/scholar?q=Amoeba%3A+Runtime+Tensor+Parallel+Transformation+for+LLM+Inference+Services 12. UCCL-EP: Portable Expert-Parallel Communication — Ziming Mao et al., 2026 https://scholar.google.com/scholar?q=UCCL-EP%3A+Portable+Expert-Parallel+Communication 13. RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training — Wei Gao et al., 2026 https://scholar.google.com/scholar?q=RollPacker%3A+Mitigating+Long-Tail+Rollouts+for+Fast%2C+Synchronous+RL+Post-Training 14. HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing — Haochen Huang et al., 2025 https://arxiv.org/abs/2509.09420 15. fMoE: Fine-Grained Expert Offloading for Large Mixture-of-Experts Serving — Hanfei Yu et al., 2025 https://arxiv.org/abs/2502.05370 16. HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference — Peng Tang et al., 2024 https://arxiv.org/abs/2411.01433 17. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu et al., 2023 https://arxiv.org/abs/2310.07240 18. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim et al., 2025 https://arxiv.org/abs/2505.23416 19. AI Post Transformers: JANUS for Scalable MoE Inference — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-15-janus-for-scalable-moe-inference-78ae30.mp3 20. AI Post Transformers: Serving MoE Models with Disaggregated Expert Parallelism — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-19-serving-moe-models-with-disaggregated-ex-6979d2.mp3 21. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3 22. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3 23. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3 24. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3

Episode metadata supplied by the publisher feed · Published Jun 29, 2026

Embed this episode

NOW PLAYING

Moebius: Seamless Parallelism Switching for MoE Serving

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 29, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!