EPISODE · Jul 29, 2026
Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
from AI Post Transformers
This episode dives into DUAL-BLADE, a systems-engineering paper examining why naive NVMe offloading of transformer KV-caches breaks down on memory-constrained edge devices. The discussion traces three compounding failures in the standard mmap-and-let-the-OS-page-cache approach: decode-phase thrashing from generic LRU eviction policies that don't understand autoregressive access patterns, prefill-phase write stalls from synchronous write-back pressure, and sequential-locality loss as the kernel's block layer fragments and reorders I/O across hardware queues. It contrasts this with FlexLLMGen's baseline approach (itself descended from Stanford's 2023 FlexGen) and explains how DUAL-BLADE's KV Placement Unit design routes tensors around these bottlenecks using cgroup-aware memory budgeting. Listeners interested in the gap between transformer-level research and the storage-stack realities of running large context windows on single-GPU edge hardware — Jetson-class devices and unified-memory workstations — will find the layer-by-layer diagnosis of kernel, block-layer, and SSD queueing behavior a rare level of systems rigor applied to an LLM-serving problem. Sources: 1. DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference — Bodon Jeong, Hongsu Byun, Youngjae Kim, Weikuan Yu, Kyungkeun Lee, Jihoon Yang, Sungyong Park, 2026 http://arxiv.org/abs/2604.26557 2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 3. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Daniel Y. Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E. Gonzalez, Percy Liang, Christopher Ré, Ion Stoica, Ce Zhang, 2023 https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU 4. LLM in a Flash: Efficient Large Language Model Inference with Limited Memory — Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, Karen Khatamifard, Minsik Cho, Carlo C. Del Mundo, Mohammad Rastegari, Mehrdad Farajtabar (Apple), 2024 https://scholar.google.com/scholar?q=LLM+in+a+Flash%3A+Efficient+Large+Language+Model+Inference+with+Limited+Memory 5. PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU — Yixin Song, Zeyu Mi, Haotong Xie, Haibo Chen (Shanghai Jiao Tong University, IPADS), 2023 https://scholar.google.com/scholar?q=PowerInfer%3A+Fast+Large+Language+Model+Serving+with+a+Consumer-grade+GPU 6. KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference — H. Zhang, C. Xia, Z. Wang, 2025 https://scholar.google.com/scholar?q=KVSwap%3A+Disk-aware+KV+Cache+Offloading+for+Long-Context+On-device+Inference 7. InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management — W. Lee, J. Lee, J. Seo, J. Sim, 2024 (OSDI '24) https://scholar.google.com/scholar?q=InfiniGen%3A+Efficient+Generative+Inference+of+Large+Language+Models+with+Dynamic+KV+Cache+Management 8. Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention (AttentionStore) — B. Gao, Z. He, P. Sharma, Q. Kang, D. Jevdjic, et al., 2024 (USENIX ATC '24) https://scholar.google.com/scholar?q=Cost-Efficient+Large+Language+Model+Serving+for+Multi-turn+Conversations+with+CachedAttention+%28AttentionStore%29 9. LMCache: An Efficient KV Cache Layer for Enterprise-scale LLM Inference — Y. Liu, Y. Cheng, J. Yao, Y. An, X. Chen, et al., 2025 https://scholar.google.com/scholar?q=LMCache%3A+An+Efficient+KV+Cache+Layer+for+Enterprise-scale+LLM+Inference 10. GPU-initiated On-demand High-throughput Storage Access in the BaM System Architecture — Z. Qureshi, V. S. Mailthody, I. Gelado, S. Min, et al., 2023 (ASPLOS '23) https://scholar.google.com/scholar?q=GPU-initiated+On-demand+High-throughput+Storage+Access+in+the+BaM+System+Architecture Interactive Visualization: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
Embed this episode
NOW PLAYING
Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
No transcript for this episode yet
Similar Episodes
No similar episodes found.