EPISODE · May 15, 2026
JANUS for Scalable MoE Inference
from AI Post Transformers
This episode explores JANUS, a systems approach to serving mixture-of-experts transformers efficiently by separating attention layers from expert layers instead of deploying the whole model as a single monolithic unit. It explains why MoE models can still be expensive and latency-prone in practice: even if only a few experts activate per token, the system must still manage large expert memory footprints, skewed expert demand, and strict token-level latency targets such as time per output token. The discussion focuses on JANUS’s core ideas, including separate GPU pools for attention and expert computation, an adaptive two-phase communication scheme that reduces cross-node messaging overhead, and SLO-aware scaling that adjusts attention and expert capacity independently. Listeners would find it interesting because it turns MoE inference from a simple “sparse compute saves money” story into a deeper argument about distributed systems design, load balancing, and the real bottlenecks that determine whether advanced models feel fast in production. Sources: 1. Janus: Disaggregating Attention and Experts for Scalable MoE Inference — Zhexiang Zhang, Ye Wang, Yumiao Zhao, Jiayu Xiao, Qianjing Yang, Xiangyu Wang, Jingzhe Jiang, Qizhen Weng, Ruichuan Chen, Shaohuai Shi, Adel N. Toosi, Yin Chen, Minchen Yu, 2025 http://arxiv.org/abs/2512.13525 2. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity — William Fedus, Barret Zoph, Noam Shazeer, 2021 https://scholar.google.com/scholar?q=Switch+Transformers%3A+Scaling+to+Trillion+Parameter+Models+with+Simple+and+Efficient+Sparsity 3. FastMoE: A Fast Mixture-of-Expert Training System — Jiaao He, Jiezhong Qiu, Aohan Zeng, Zhilin Yang, Jidong Zhai, Jie Tang, 2021 https://scholar.google.com/scholar?q=FastMoE%3A+A+Fast+Mixture-of-Expert+Training+System 4. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024 https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving 5. MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism — Ruidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu, Cesar A. Stuardo, Dongyang Wang and others, 2025 https://scholar.google.com/scholar?q=MegaScale-Infer%3A+Serving+Mixture-of-Experts+at+Scale+with+Disaggregated+Expert+Parallelism 6. eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference — Suraiya Tairin, Shohaib Mahmud, Haiying Shen, Anand Iyer, 2025 https://scholar.google.com/scholar?q=eMoE%3A+Task-aware+Memory+Efficient+Mixture-of-Experts-Based+%28MoE%29+Model+Inference 7. SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference — Luchang Li, Dongfang Li, Bozhao Gong, Yu Zhang, 2026 https://scholar.google.com/scholar?q=SLO-Aware+Compute+Resource+Allocation+for+Prefill-Decode+Disaggregated+LLM+Inference 8. MoEless: Efficient MoE LLM Serving via Serverless Computing — Hanfei Yu, Bei Ouyang, Shwai He, Ang Li, Hao Wang, 2026 https://scholar.google.com/scholar?q=MoEless%3A+Efficient+MoE+LLM+Serving+via+Serverless+Computing 9. xDeepServe: Model-as-a-Service on Huawei CloudMatrix384 — Ao Xiao et al., 2025 https://scholar.google.com/scholar?q=xDeepServe%3A+Model-as-a-Service+on+Huawei+CloudMatrix384 10. Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling — Yan Li, Zhenyu Zhang, Zhengang Wang, Pengfei Chen, Pengfei Zheng, 2025 https://scholar.google.com/scholar?q=Semantic+Parallelism%3A+Redefining+Efficient+MoE+Inference+via+Model-Data+Co-Scheduling 11. GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference — Yu Han, Lehan Pan, Jie Peng, Ziyang Tao, Hanqi Zhu, Wuyang Zhang, Yanyong Zhang, 2025 https://scholar.google.com/scholar?q=GRACE-MoE%3A+Grouping+and+Replication+with+Locality-Aware+Routing+for+Efficient+Distributed+MoE+Inference 12. BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems — Yuxin Wang, Yuhan Chen, Zeyu Li, Xueze Kang, Yuchu Fang, Yeju Zhou, Yang Zheng, Zhenheng Tang, Xin He, Rui Guo, Xin Wang, Qiang Wang, Amelie Chi Zhou, Xiaowen Chu, 2024 https://scholar.google.com/scholar?q=BurstGPT%3A+A+Real-world+Workload+Dataset+to+Optimize+LLM+Serving+Systems 13. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI et al., 2024 https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model 14. Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference — Ranggi Hwang et al., 2023/2024 https://arxiv.org/abs/2308.12066 15. HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference — Peng Tang et al., 2024 https://arxiv.org/abs/2411.01433 16. DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference — Yujie Zhang, Shivam Aggarwal, Tulika Mitra, 2025 https://arxiv.org/abs/2501.10375 17. AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding — Zikun Li et al., 2025 https://arxiv.org/abs/2501.12162 18. SLOs-Serve: Optimized Serving of Multi-SLO LLMs — Siyuan Chen et al., 2025 https://arxiv.org/abs/2504.08784 19. Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement — Tian Wu et al., 2025 https://arxiv.org/abs/2508.12851 20. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3 21. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3 22. AI Post Transformers: Computation-Bandwidth-Memory Trade-offs for AI Infrastructure — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-09-computation-bandwidth-memory-trade-offs-a83f2b.mp3 23. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3 24. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3 25. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3 26. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3 Interactive Visualization: JANUS for Scalable MoE Inference
Embed this episode
NOW PLAYING
JANUS for Scalable MoE Inference
No transcript for this episode yet
Similar Episodes
No similar episodes found.