ASAP: Disaggregating Attention and Experts for Faster MoE Prefill episode artwork

EPISODE · Jul 26, 2026

ASAP: Disaggregating Attention and Experts for Faster MoE Prefill

from AI Post Transformers

This episode explores ASAP, a Huawei-developed serving system for mixture-of-experts models that physically separates the attention and expert computation stages onto different hardware with non-blocking communication between them. The discussion traces the diagnosis behind the design: production traces reveal a "straggler effect" where global synchronization barriers between attention's data-parallel groups and the shared expert pool force every group to wait on the slowest one, and this imbalance is mathematically unavoidable since attention cost scales with the sum of squares of sequence lengths rather than total tokens. The hosts draw a parallel to the classic MapReduce straggler problem, framing this as a structural consequence of combining Expert Parallelism for MoE layers with Data Parallelism for attention rather than something a better scheduler could fix. Listeners get a grounded walkthrough of prefill versus decode, time-to-first-token, and why hybrid parallelism setups like DeepSeek-V3's TP=8/DP=4/EP=32 configuration create this bottleneck in the first place, setting up the paper's disaggregated architecture as a direct response to a precisely characterized failure mode rather than a speculative fix. Sources: 1. ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill — Weiwei Chen, Shuang Chen, Lele Li, Qiang Hu, Han Li, Xin Ye, Ming Yan, Zhibin Yu, 2026 http://arxiv.org/abs/2606.22541 2. Revealing the challenges of attention-ffn disaggregation for modern MoE models and hardware systems — Guowei Liu, Hongming Li, Yaning Guo, Yongxi Lyu, Mo Zhou, Yi Liu, Zhaogeng Li, Yanpeng Wang, 2026 https://scholar.google.com/scholar?q=Revealing+the+challenges+of+attention-ffn+disaggregation+for+modern+MoE+models+and+hardware+systems 3. Expert-as-a-Service: Towards efficient, scalable, and robust large-scale MoE serving — Ziming Liu, Boyu Tian, Guoteng Wang, Zhen Jiang, et al., 2025 https://scholar.google.com/scholar?q=Expert-as-a-Service%3A+Towards+efficient%2C+scalable%2C+and+robust+large-scale+MoE+serving 4. MegaScale-Infer — Cited as [55] in ASAP, 2025/2026 https://scholar.google.com/scholar?q=MegaScale-Infer 5. Step-3 — Cited as [44] in ASAP, 2025/2026 https://scholar.google.com/scholar?q=Step-3 6. Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbots — Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2025 https://scholar.google.com/scholar?q=Mooncake%3A+Trading+more+storage+for+less+computation+%E2%80%94+a+KVCache-centric+architecture+for+serving+LLM+chatbots 7. Sarathi: Efficient LLM inference by piggybacking decodes with chunked prefills — Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S Gulavani, Ramachandran Ramjee, 2023 https://scholar.google.com/scholar?q=Sarathi%3A+Efficient+LLM+inference+by+piggybacking+decodes+with+chunked+prefills 8. TokenWeave: Efficient compute-communication overlap for distributed LLM inference — Raja Gond, Nipun Kwatra, Ramachandran Ramjee, 2025 https://scholar.google.com/scholar?q=TokenWeave%3A+Efficient+compute-communication+overlap+for+distributed+LLM+inference Interactive Visualization: ASAP: Disaggregating Attention and Experts for Faster MoE Prefill

Episode metadata supplied by the publisher feed · Published Jul 26, 2026

Embed this episode

NOW PLAYING

ASAP: Disaggregating Attention and Experts for Faster MoE Prefill

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on July 26, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!