Harvest: Borrowing Peer GPU Memory for LLMs episode artwork

EPISODE · Jun 5, 2026

Harvest: Borrowing Peer GPU Memory for LLMs

from AI Post Transformers

This episode explores Harvest, a system for LLM inference that uses idle HBM on neighboring NVLink-connected GPUs as a temporary cache when a serving GPU runs out of local memory. It explains why LLM serving is often bottlenecked more by memory capacity and data movement than by raw compute, focusing on two concrete cases: growing KV caches during long-context decoding and the shifting expert weights used in mixture-of-experts models. A key argument is that peer GPU memory is only useful if it is revocable without breaking correctness, so Harvest treats borrowed memory as a best-effort cache backed by authoritative copies or reconstruction paths elsewhere. Listeners get specific performance results, including up to 5.65x lower KV-cache transfer latency than CPU offload, 7.5x to 9.5x faster expert transfers over NVLink, and roughly 1.5x to 2.0x throughput gains on models such as Qwen and Phi-3.5. Sources: 1. Harvest: Opportunistic Peer-to-Peer GPU Caching for LLM Inference — Nikhil Gopal, Kostis Kaffes, 2026 http://arxiv.org/abs/2602.00328 2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 3. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Beidi Chen, Percy Liang, Christopher Re, Ion Stoica, Ce Zhang, 2023 https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU 4. Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models — Keisuke Kamahori, Tian Tang, Yile Gu, Kan Zhu, Baris Kasikci, 2025 https://scholar.google.com/scholar?q=Fiddler%3A+CPU-GPU+Orchestration+for+Fast+Inference+of+Mixture-of-Experts+Models 5. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, et al., 2025 https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving 6. AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains — Abhishek Vijaya Kumar, Gianni Antichi, Rachee Singh, 2025 https://scholar.google.com/scholar?q=AQUA%3A+Network-Accelerated+Memory+Offloading+for+LLMs+in+Scale-Up+GPU+Domains 7. Accurate Expert Predictions in MoE Inference via Cross-Layer Gate — Zhiyuan Fang, Hong Huang, Yiming Lyu, Jiyang Chen, Yu Yu, Zexi Zheng, 2025 https://scholar.google.com/scholar?q=Accurate+Expert+Predictions+in+MoE+Inference+via+Cross-Layer+Gate 8. Characterization of Large Language Model Development in the Datacenter — Qinghao Hu et al., 2024 https://scholar.google.com/scholar?q=Characterization+of+Large+Language+Model+Development+in+the+Datacenter 9. MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters — Qizhen Weng et al., 2022 https://scholar.google.com/scholar?q=MLaaS+in+the+Wild%3A+Workload+Analysis+and+Scheduling+in+Large-Scale+Heterogeneous+GPU+Clusters 10. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu et al., 2024 https://arxiv.org/abs/2406.17565 11. KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows — Zaifeng Pan et al., 2025 https://arxiv.org/abs/2507.07400 12. Learned Prefix Caching for Efficient LLM Inference — Dongsheng Yang et al., 2025 https://papers.neurips.cc/paper_files/paper/2025/hash/414f642a1ea9350006669774cba9bcd4-Abstract-Conference.html 13. DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance — Yuning Zhang et al., 2025 https://arxiv.org/abs/2509.07379 14. ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference — Zixu Shen et al., 2025 https://arxiv.org/abs/2510.26730 15. AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference — Shuzhang Zhong et al., 2024 https://arxiv.org/abs/2408.10284 16. HyperAttention: Long-context Attention in Near-Linear Time — Insu Han et al., 2023 https://arxiv.org/abs/2310.05869 17. Beyond Attention: Breaking the Limits of Transformer Context Length with Recurrent Memory — Aydar Bulatov et al., 2024 https://ojs.aaai.org/index.php/AAAI/article/download/29722/31239 18. AI Post Transformers: ScoutAttention for Efficient KV Cache Offloading — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-24-scoutattention-for-efficient-kv-cache-of-b26699.mp3 19. AI Post Transformers: JANUS for Scalable MoE Inference — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-15-janus-for-scalable-moe-inference-78ae30.mp3 20. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3 21. AI Post Transformers: FlexGen: High-Throughput LLM Inference on a Single GPU — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/flexgen-high-throughput-llm-inference-on-a-single-gpu/ 22. AI Post Transformers: Stochastic KV Routing for Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-29-stochastic-kv-routing-for-cache-sharing-5fef63.mp3 23. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3

Episode metadata supplied by the publisher feed · Published Jun 5, 2026

Embed this episode

NOW PLAYING

Harvest: Borrowing Peer GPU Memory for LLMs

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 5, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!