EPISODE · Jul 26, 2026
ASAP: Disaggregating Attention and Experts for Faster MoE Prefill
from AI Post Transformers
This episode explores ASAP, a Huawei-developed serving system for mixture-of-experts models that physically separates the attention and expert computation stages onto different hardware with non-blocking communication between them. The discussion traces the diagnosis behind the design: production traces reveal a "straggler effect" where global synchronization barriers between attention's data-parallel groups and the shared expert pool force every group to wait on the slowest one, and this imbalance is mathematically unavoidable since attention cost scales with the sum of squares of sequence lengths rather than total tokens. The hosts draw a parallel to the classic MapReduce straggler problem, framing this as a structural consequence of combining Expert Parallelism for MoE layers with Data Parallelism for attention rather than something a better scheduler could fix. Listeners get a grounded walkthrough of prefill versus decode, time-to-first-token, and why hybrid parallelism setups like DeepSeek-V3's TP=8/DP=4/EP=32 configuration create this bottleneck in the first place, setting up the paper's disaggregated architecture as a direct response to a precisely characterized failure mode rather than a speculative fix. Sources: 1. ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill — Weiwei Chen, Shuang Chen, Lele Li, Qiang Hu, Han Li, Xin Ye, Ming Yan, Zhibin Yu, 2026 http://arxiv.org/abs/2606.22541 2. Revealing the challenges of attention-ffn disaggregation for modern MoE models and hardware systems — Guowei Liu, Hongming Li, Yaning Guo, Yongxi Lyu, Mo Zhou, Yi Liu, Zhaogeng Li, Yanpeng Wang, 2026 https://scholar.google.com/scholar?q=Revealing+the+challenges+of+attention-ffn+disaggregation+for+modern+MoE+models+and+hardware+systems 3. Expert-as-a-Service: Towards efficient, scalable, and robust large-scale MoE serving — Ziming Liu, Boyu Tian, Guoteng Wang, Zhen Jiang, et al., 2025 https://scholar.google.com/scholar?q=Expert-as-a-Service%3A+Towards+efficient%2C+scalable%2C+and+robust+large-scale+MoE+serving 4. MegaScale-Infer — Cited as [55] in ASAP, 2025/2026 https://scholar.google.com/scholar?q=MegaScale-Infer 5. Step-3 — Cited as [44] in ASAP, 2025/2026 https://scholar.google.com/scholar?q=Step-3 6. Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbots — Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2025 https://scholar.google.com/scholar?q=Mooncake%3A+Trading+more+storage+for+less+computation+%E2%80%94+a+KVCache-centric+architecture+for+serving+LLM+chatbots 7. Sarathi: Efficient LLM inference by piggybacking decodes with chunked prefills — Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S Gulavani, Ramachandran Ramjee, 2023 https://scholar.google.com/scholar?q=Sarathi%3A+Efficient+LLM+inference+by+piggybacking+decodes+with+chunked+prefills 8. TokenWeave: Efficient compute-communication overlap for distributed LLM inference — Raja Gond, Nipun Kwatra, Ramachandran Ramjee, 2025 https://scholar.google.com/scholar?q=TokenWeave%3A+Efficient+compute-communication+overlap+for+distributed+LLM+inference Interactive Visualization: ASAP: Disaggregating Attention and Experts for Faster MoE Prefill
Embed this episode
NOW PLAYING
ASAP: Disaggregating Attention and Experts for Faster MoE Prefill
No transcript for this episode yet
Similar Episodes
No similar episodes found.