EPISODE · Jun 24, 2026
AI+HW 2035: Co-Designing Efficient AI Systems
from AI Post Transformers
This episode explores the AI+HW 2035 roadmap, arguing that the next decade of AI progress will depend less on raw compute growth and more on coordinated design across models, compilers, runtimes, memory systems, and chips. It breaks down the memory wall in concrete terms, showing how moving weights, activations, and KV caches can cost more time and energy than the math itself, especially for inference, autoregressive serving, and state-heavy workloads like video world models. The discussion examines quantization, mixed precision, sparsity, pruning, distillation, tiered memory, IO-aware attention, and hardware-aware scheduling, with the key claim that these methods only matter when the full stack preserves locality and avoids wasteful data movement. Listeners would find it interesting because it treats AI efficiency as a practical systems problem and policy agenda, not just a matter of inventing better model architectures. Sources: 1. AI+HW 2035: Shaping the Next Decade — Deming Chen, Jason Cong, Azalia Mirhoseini, Christos Kozyrakis, Subhasish Mitra, Jinjun Xiong, Cliff Young, Anima Anandkumar, Michael Littman, Aron Kirschen, Sophia Shao, Serge Leef, Naresh Shanbhag, Dejan Milojicic, Michael Schulte, Gert Cauwenberghs, Jerry M. Chow, Tri Dao, Kailash Gopalakrishnan, Richard Ho, Hoshik Kim, Kunle Olukotun, David Z. Pan, Mark Ren, Dan Roth, Aarti Singh, Yizhou Sun, Yusu Wang, Yann LeCun, Ruchir Puri, 2026 http://arxiv.org/abs/2603.05225 2. Mixed Precision Training — Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, and others, 2017 https://scholar.google.com/scholar?q=Mixed+Precision+Training 3. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference — Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, and others, 2018 https://scholar.google.com/scholar?q=Quantization+and+Training+of+Neural+Networks+for+Efficient+Integer-Arithmetic-Only+Inference 4. FP8 Formats for Deep Learning — Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, and others, 2022 https://scholar.google.com/scholar?q=FP8+Formats+for+Deep+Learning 5. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Chuang Gan, Song Han, and others, 2023 https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration 6. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao et al., 2022 https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness 7. A Full-Stack Search Technique for Domain Optimized Deep Learning Accelerators — Dan Zhang et al., 2022 https://scholar.google.com/scholar?q=A+Full-Stack+Search+Technique+for+Domain+Optimized+Deep+Learning+Accelerators 8. A Compute-in-Memory Chip Based on Resistive Random-Access Memory — Weier Wan et al., 2022 https://scholar.google.com/scholar?q=A+Compute-in-Memory+Chip+Based+on+Resistive+Random-Access+Memory 9. Intelligence per Watt: Measuring Intelligence Efficiency of Local AI — Jon Saad-Falcon et al., 2025 https://scholar.google.com/scholar?q=Intelligence+per+Watt%3A+Measuring+Intelligence+Efficiency+of+Local+AI 10. PinDrop: Breaking the Silence on SDCs in a Large-Scale Fleet — Peter W. Deutsch et al., 2026 https://scholar.google.com/scholar?q=PinDrop%3A+Breaking+the+Silence+on+SDCs+in+a+Large-Scale+Fleet 11. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yuhan Liu et al., 2025 https://scholar.google.com/scholar?q=LMCache%3A+An+Efficient+KV+Cache+Layer+for+Enterprise-Scale+LLM+Inference 12. TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale — Dongha Yoon et al., 2025 https://scholar.google.com/scholar?q=TraCT%3A+Disaggregated+LLM+Serving+with+CXL+Shared+Memory+KV+Cache+at+Rack-Scale 13. Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads — Cunchen Hu et al., 2024 https://scholar.google.com/scholar?q=Inference+without+Interference%3A+Disaggregate+LLM+Inference+for+Mixed+Downstream+Workloads 14. CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving — Dong Liu and Yanxuan Yu, 2025 https://scholar.google.com/scholar?q=CXL-SpecKV%3A+A+Disaggregated+FPGA+Speculative+KV-Cache+for+Datacenter+LLM+Serving 15. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025 https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse 16. Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference — Kexin Chu et al., 2025 https://scholar.google.com/scholar?q=Selective+KV-Cache+Sharing+to+Mitigate+Timing+Side-Channels+in+LLM+Inference 17. Test-Time Training on Nearest Neighbors for Large Language Models — Moritz Hardt and Yu Sun, 2023 https://scholar.google.com/scholar?q=Test-Time+Training+on+Nearest+Neighbors+for+Large+Language+Models 18. Test-Time Learning for Large Language Models — Jinwu Hu et al., 2025 https://scholar.google.com/scholar?q=Test-Time+Learning+for+Large+Language+Models 19. AI Post Transformers: FlatAttention for Tile-Based Accelerator Inference — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-04-flatattention-for-tile-based-accelerator-56e6ca.mp3 20. AI Post Transformers: LPU Chip for Low-Latency LLM Inference — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-20-lpu-chip-for-low-latency-llm-inference-be13c3.mp3 21. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3 22. AI Post Transformers: Mistral 7B: Superior Performance in a Smaller Package — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/mistral-7b-superior-performance-in-a-smaller-package/ 23. AI Post Transformers: PALOMA: Benchmarking Language Model Fit Across Domains — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-23-paloma-benchmarking-language-model-fit-a-360060.mp3 24. AI Post Transformers: Automating CNN Mapping on Embedded FPGAs — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-06-automating-cnn-mapping-on-embedded-fpgas-4c1dc3.mp3 25. AI Post Transformers: FPGA Neural Network Accelerators for Space — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-26-fpga-neural-network-accelerators-for-spa-3087ae.mp3 26. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
Embed this episode
NOW PLAYING
AI+HW 2035: Co-Designing Efficient AI Systems
No transcript for this episode yet
Similar Episodes
No similar episodes found.