EPISODE · Aug 10, 2026
Move the Query, Not the Cache: MLA Rewrites GPU Fabric Attention Routing
from AI Post Transformers
This episode explores cross-instance attention in disaggregated LLM serving, focusing on the surprising size inversion created by Multi-head Latent Attention: a routed decoding query shrinks to roughly a kilobyte while the cache chunk it must read can balloon to 61 megabytes across layers, upending the old assumption that query and cache are comparably sized. The discussion traces why this scenario is becoming routine — providers sharing precomputed caches for large corpora that outgrow a single GPU's memory, and agentic workloads where many sub-agents query one oversized shared prefix — and lays out the three possible strategies (route, fetch, or recompute locally) for handling the mismatch, including how sparse indexers further shrink the routable unit to scattered top-k blocks. A key thread examines device-initiated RDMA via IBGDA, challenging the intuition that skipping the CPU proxy is automatically faster: prior work on tiny mixture-of-experts messages actually found IBGDA slower, but the paper's controlled test on kilobyte-scale attention traffic shows the CPU-proxy path is 40% slower at the median and over 50% slower at steady state. Listeners interested in GPU networking, KV-cache architecture, or the practical plumbing behind large-scale LLM inference will find the paper's empirical resolution of a previously untested assumption particularly compelling. Sources: 1. Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics — Bole Ma, Jan Eitzinger, Harald Köstler, Gerhard Wellein, 2026 http://arxiv.org/abs/2606.01502 2. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI (research team), 2024 https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model 3. DeepSeek-V3 Technical Report — DeepSeek-AI (research team), 2024 https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report 4. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, Sumit Sanghai, 2023 https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints 5. TransMLA: Multi-Head Latent Attention Is All You Need — Fanxu Meng, Zengwei Yao, Muhan Zhang, et al., 2025 https://scholar.google.com/scholar?q=TransMLA%3A+Multi-Head+Latent+Attention+Is+All+You+Need 6. Improving Network Performance of HPC Systems Using NVIDIA Magnum IO NVSHMEM and GPUDirect Async — NVIDIA (NVSHMEM / Magnum IO engineering team), 2023 https://scholar.google.com/scholar?q=Improving+Network+Performance+of+HPC+Systems+Using+NVIDIA+Magnum+IO+NVSHMEM+and+GPUDirect+Async 7. Efficient Inter-node MPI Communication using GPUDirect RDMA for InfiniBand Clusters with NVIDIA GPUs — Sreeram Potluri, Khaled Hamidouche, Akshay Venkatesh, Devendar Bureddy, Dhabaleswar K. Panda, 2013 https://scholar.google.com/scholar?q=Efficient+Inter-node+MPI+Communication+using+GPUDirect+RDMA+for+InfiniBand+Clusters+with+NVIDIA+GPUs 8. DeepEP: an efficient expert-parallel communication library (and related DeepSeek-V3 Technical Report communication sections) — DeepSeek-AI (research/infra team), 2025 / 2024 https://scholar.google.com/scholar?q=DeepEP%3A+an+efficient+expert-parallel+communication+library+%28and+related+DeepSeek-V3+Technical+Report+communication+sections%29 9. Introducing OpenSHMEM: SHMEM for the PGAS Community — Barbara Chapman, Tony Curtis, Swaroop Pophale, et al., 2010 https://scholar.google.com/scholar?q=Introducing+OpenSHMEM%3A+SHMEM+for+the+PGAS+Community 10. Preble: Efficient Distributed Prompt Scheduling for LLM Serving — Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, Yiying Zhang, 2024 https://scholar.google.com/scholar?q=Preble%3A+Efficient+Distributed+Prompt+Scheduling+for+LLM+Serving 11. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu, Heyang Huang, Junhao Hu, Jiang Xu, Xusheng Chen, Tao Xie, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, Yizhou Shan, 2024 https://scholar.google.com/scholar?q=MemServe%3A+Context+Caching+for+Disaggregated+LLM+Serving+with+Elastic+Memory+Pool 12. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, Beidi Chen, 2023 https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models Interactive Visualization: Move the Query, Not the Cache: MLA Rewrites GPU Fabric Attention Routing
Embed this episode
NOW PLAYING
Move the Query, Not the Cache: MLA Rewrites GPU Fabric Attention Routing
No transcript for this episode yet
Similar Episodes
No similar episodes found.