EMO: Emergent Modularity for Mixture-of-Experts episode artwork

EPISODE · Jun 22, 2026

EMO: Emergent Modularity for Mixture-of-Experts

from AI Post Transformers

Hal Turing and Dr. Ada Shannon take a deep dive into EMO: Pretraining Mixture of Experts for Emergent Modularity, a May 7, 2026 paper by Ryan Wang and co-authors from UC Berkeley and the Allen Institute for AI. The episode centers on a practical deployment question: if a workload is mostly code, math, or biomed, why must operators keep an entire giant model in memory instead of loading only the relevant slice? They frame EMO against the broader rise of sparse Mixture-of-Experts systems and explain why industry progress on active-parameter efficiency is not the same as delivering clean, domain-specific modules that can stand on their own at inference time. The discussion carefully separates standard MoE behavior from the stronger notion of modularity that EMO is targeting. Hal and Ada walk through how sparse-gated MoE and Switch Transformer style routing already allow different tokens to activate different experts, but argue that this still leaves deployment looking monolithic because the router makes local token-level decisions rather than exposing stable task-level components. A biology prompt can still scatter across a messy set of experts, and the next sentence may hit a different set entirely. The hosts use that distinction to unpack the paper’s core concepts: emergent modularity from unlabeled data, semantic expert specialization around meaningful domains like code or math, composable architecture, and the memory-accuracy frontier that determines whether smaller loaded expert pools can preserve real capability. The episode then gets into EMO’s training design and why the method is more than a single routing tweak. Ada explains the paper’s two-level routing scheme, where a document first selects a shared candidate pool of experts and individual tokens then choose active experts only within that pool, forcing document-consistent structure without removing all local flexibility. They also cover the supporting recipe: random pool-size sampling to expose the model to different memory budgets during training, global load balancing so a few experts do not dominate usage, and document-length-aware training so very long documents do not overwhelm the learning signal. The result is a focused discussion of whether MoE pretraining can produce expert groups that are not just sparsely activated, but genuinely deployable as modular tools. Sources: 1. EMO: Emergent Modularity for Mixture-of-Experts https://allenai.org/papers/emo 2. DEMix Layers: Disentangling Domains for Modular Language Modeling — Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A. Smith, Luke Zettlemoyer, 2021 https://scholar.google.com/scholar?q=DEMix+Layers%3A+Disentangling+Domains+for+Modular+Language+Modeling 3. ModuleFormer: Modularity Emerges from Mixture-of-Experts — Yikang Shen, Zheyu Zhang, Tianyou Cao, Shawn Tan, Zhenfang Chen, Chuang Gan, 2023 https://scholar.google.com/scholar?q=ModuleFormer%3A+Modularity+Emerges+from+Mixture-of-Experts 4. OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models — Fuzhao Xue, Zian Zheng, Yao Fu, Jinjie Ni, Zangwei Zheng, Wangchunshu Zhou, Yang You, 2024 https://scholar.google.com/scholar?q=OpenMoE%3A+An+Early+Effort+on+Open+Mixture-of-Experts+Language+Models 5. EMO: Pretraining Mixture of Experts for Emergent Modularity — Ryan Wang, Akshita Bhagia, Sewon Min, 2026 https://scholar.google.com/scholar?q=EMO%3A+Pretraining+Mixture+of+Experts+for+Emergent+Modularity 6. FlexOlmo: Open Language Models for Flexible Data Use — Weijia Shi et al., 2025 https://scholar.google.com/scholar?q=FlexOlmo%3A+Open+Language+Models+for+Flexible+Data+Use 7. Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM — Sainbayar Sukhbaatar et al., 2024 https://scholar.google.com/scholar?q=Branch-Train-MiX%3A+Mixing+Expert+LLMs+into+a+Mixture-of-Experts+LLM 8. The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise — Xi Wang, Soufiane Hayou, Eric Nalisnick, 2026 https://scholar.google.com/scholar?q=The+Myth+of+Expert+Specialization+in+MoEs%3A+Why+Routing+Reflects+Geometry%2C+Not+Necessarily+Domain+Expertise 9. Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations — Zican Dong et al., 2025 https://scholar.google.com/scholar?q=Domain-Specific+Pruning+of+Large+Mixture-of-Experts+Models+with+Few-shot+Demonstrations 10. Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs — Enshu Liu et al., 2024 https://scholar.google.com/scholar?q=Efficient+Expert+Pruning+for+Sparse+Mixture-of-Experts+Language+Models%3A+Enhancing+Performance+and+Reducing+Inference+Costs 11. MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router — Yanyue Xie et al., 2024 https://scholar.google.com/scholar?q=MoE-Pruner%3A+Pruning+Mixture-of-Experts+Large+Language+Model+using+the+Hints+from+Its+Router 12. AdaMoE: Token-Adaptive Routing with Null Experts for Mixture-of-Experts Language Models — Zihao Zeng et al., 2024 https://scholar.google.com/scholar?q=AdaMoE%3A+Token-Adaptive+Routing+with+Null+Experts+for+Mixture-of-Experts+Language+Models 13. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models — Damai Dai et al., 2024 https://scholar.google.com/scholar?q=DeepSeekMoE%3A+Towards+Ultimate+Expert+Specialization+in+Mixture-of-Experts+Language+Models 14. Fast Inference of Mixture-of-Experts Language Models with Offloading — Artyom Eliseev and Denis Mazur, 2023 https://scholar.google.com/scholar?q=Fast+Inference+of+Mixture-of-Experts+Language+Models+with+Offloading 15. Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding — Zhibin Wang et al., 2025 https://scholar.google.com/scholar?q=Accelerating+Mixture-of-Experts+Inference+by+Hiding+Offloading+Latency+with+Speculative+Decoding 16. AI Post Transformers: EMO: Emergent Modularity in Sparse Language Models — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-06-emo-emergent-modularity-in-sparse-langua-9551c4.mp3 17. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3 18. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3 19. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3 Interactive Visualization: EMO: Emergent Modularity for Mixture-of-Experts

Episode metadata supplied by the publisher feed · Published Jun 22, 2026

Embed this episode

NOW PLAYING

EMO: Emergent Modularity for Mixture-of-Experts

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 22, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!