PODCAST · technology
AI Post Transformers
by mcgrof
AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.
-
708
Reptile: The First-Order Meta-Learning Shortcut That Works
This episode examines "On First-Order Meta-Learning Algorithms" by Alex Nichol, Joshua Achiam, and John Schulman, which challenges the assumption that MAML's expensive second-derivative computation is essential for effective few-shot learning. The discussion traces the lineage from MAML's nested optimization — where an outer loop backpropagates through an inner loop's gradient steps via the Hessian — through First-Order MAML's approximation, to Reptile, a stripped-down algorithm that simply runs SGD on sampled tasks and nudges the initialization toward the result, with no meta-gradient or train-test split required. A central tension drives the conversation: why pulling an initialization toward "wherever SGD landed" produces a genuinely different target than plain joint training across tasks, rather than just averaging into one generic model. The hosts set up a Taylor-expansion argument to explain which gradient terms MAML, FOMAML, and Reptile weight differently, revealing the mathematical reason the cheaper approximation retains nearly all the useful signal. Listeners interested in the mechanics of meta-learning, gradient-based optimization tradeoffs, or the history of few-shot learning approaches will find the paper's practical implications for scaling meta-learning algorithms especially relevant. Sources: 1. On First-Order Meta-Learning Algorithms — Alex Nichol, Joshua Achiam, John Schulman, 2018 http://arxiv.org/abs/1803.02999 2. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks — Chelsea Finn, Pieter Abbeel, Sergey Levine, 2017 https://scholar.google.com/scholar?q=Model-Agnostic+Meta-Learning+for+Fast+Adaptation+of+Deep+Networks 3. How to train your MAML — Antreas Antoniou, Harrison Edwards, Amos Storkey, 2019 https://scholar.google.com/scholar?q=How+to+train+your+MAML 4. Meta-Learning with Implicit Gradients — Aravind Rajeswaran, Chelsea Finn, Sham Kakade, Sergey Levine, 2019 https://scholar.google.com/scholar?q=Meta-Learning+with+Implicit+Gradients 5. Optimization as a Model for Few-Shot Learning — Sachin Ravi, Hugo Larochelle, 2017 https://scholar.google.com/scholar?q=Optimization+as+a+Model+for+Few-Shot+Learning 6. Learning to Learn by Gradient Descent by Gradient Descent — Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W. Hoffman, David Pfau, Tom Schaul, Nando de Freitas, 2016 https://scholar.google.com/scholar?q=Learning+to+Learn+by+Gradient+Descent+by+Gradient+Descent 7. Using Fast Weights to Deblur Old Memories — Geoffrey E. Hinton, David C. Plaut, 1987 https://scholar.google.com/scholar?q=Using+Fast+Weights+to+Deblur+Old+Memories 8. Parallelized Stochastic Gradient Descent — Martin Zinkevich, Markus Weimer, Lihong Li, Alex J. Smola, 2010 https://scholar.google.com/scholar?q=Parallelized+Stochastic+Gradient+Descent
-
707
FLINT: Sharing High Bandwidth Flash for Scalable LLM Inference
This episode examines FLINT, a proposed hardware/software architecture from Huawei's Zurich research lab (with ETH Zürich and HUST) for closing the gap between multi-terabyte LLM weight sizes and the far smaller on-package memory of today's GPUs. It introduces high bandwidth flash (HBF), an emerging memory tier that stacks 3D NAND dies with through-silicon vias to sit directly beside HBM in the accelerator package, storing read-only model weights while HBM handles fast-changing KV cache and activations. The discussion walks through core NAND flash mechanics — dies, planes, blocks, and pages, along with the punishing asymmetry between microsecond reads and millisecond erases — to explain why naive flash designs stall under refresh operations and static prefetching. It then details how FLINT's burst-buffer controller replaces compiler-driven prefetch hints with real-time demand-based read coalescing, using HBF's built-in page and cache buffers instead of dedicated SRAM. Listeners interested in memory system design, inference hardware economics, and the practical engineering trade-offs of scaling capacity without wasting GPU compute will find this a detailed look at a genuinely emerging technology rather than a shipping product. Sources: 1. FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration — Geraldo F. Oliveira, Arash Tavakkol, Xiangyu Zhu, Ahmet Caner Yüzügüler, Vamanan Arulchelvan, Lukas Cavigelli, Renzo Andri, Mohammad Sadrosadati, Jia Xinglei, Onur Mutlu, Zhou Ke, Shai Bergman, Ji Zhang, 2026 http://arxiv.org/abs/2608.25062 2. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, Yuxiong He (Microsoft), 2021 https://scholar.google.com/scholar?q=ZeRO-Infinity%3A+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning 3. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, et al., 2023 https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU 4. Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System — Hongsun Jang, Jaeyong Song, Jaewon Jung, Jaeyoung Park, Youngsok Kim, Jinho Lee, 2024, HPCA https://scholar.google.com/scholar?q=Smart-Infinity%3A+Fast+Large+Language+Model+Training+using+Near-Storage+Processing+on+a+Real+System 5. InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference — Xiurui Pan, Endian Li, Qiao Li, Shengwen Liang, Yizhou Shan, Ke Zhou, Yingwei Luo, Xiaolin Wang, Jie Zhang, 2024 https://scholar.google.com/scholar?q=InstInfer%3A+In-Storage+Attention+Offloading+for+Cost-Effective+Long-Context+LLM+Inference 6. DFTL: A Flash Translation Layer Employing Demand-based Selective Caching of Page-level Address Mappings — Aayush Gupta, Youngjae Kim, Bhuvan Urgaonkar, 2009, ASPLOS https://scholar.google.com/scholar?q=DFTL%3A+A+Flash+Translation+Layer+Employing+Demand-based+Selective+Caching+of+Page-level+Address+Mappings 7. Design Tradeoffs for SSD Performance — Nitin Agrawal, Vijayan Prabhakaran, Ted Wobber, John D. Davis, Mark Manasse, Rina Panigrahy, 2008, USENIX ATC https://scholar.google.com/scholar?q=Design+Tradeoffs+for+SSD+Performance 8. ZNS: Avoiding the Block Interface Tax for Flash-based SSDs — Matias Bjørling, Abutalib Aghayev, Hans Holmberg, Aravind Ramesh, Damien Le Moal, Gregory R. Ganger, George Amvrosiadis, 2021, USENIX ATC https://scholar.google.com/scholar?q=ZNS%3A+Avoiding+the+Block+Interface+Tax+for+Flash-based+SSDs 9. RAIDR: Retention-Aware Intelligent DRAM Refresh — Jamie Liu, Ben Jaiyen, Richard Veras, Onur Mutlu, 2012, ISCA https://scholar.google.com/scholar?q=RAIDR%3A+Retention-Aware+Intelligent+DRAM+Refresh 10. Threshold Voltage Distribution in MLC NAND Flash Memory: Characterization, Analysis, and Modeling — Yu Cai, Erich F. Haratsch, Onur Mutlu, Ken Mai, 2013, DATE https://scholar.google.com/scholar?q=Threshold+Voltage+Distribution+in+MLC+NAND+Flash+Memory%3A+Characterization%2C+Analysis%2C+and+Modeling 11. Error Analysis and Retention-Aware Error Management for NAND Flash Memory — Yu Cai, Erich F. Haratsch, Onur Mutlu, Ken Mai, 2013, Intel Technology Journal https://scholar.google.com/scholar?q=Error+Analysis+and+Retention-Aware+Error+Management+for+NAND+Flash+Memory 12. H3: Hybrid Architecture Using High Bandwidth Memory and High Bandwidth Flash for Cost-Efficient LLM Inference — M. Ha, E. Kim, H. Kim, 2026 https://scholar.google.com/scholar?q=H3%3A+Hybrid+Architecture+Using+High+Bandwidth+Memory+and+High+Bandwidth+Flash+for+Cost-Efficient+LLM+Inference 13. LLM in a Flash: Efficient Large Language Model Inference with Limited Memory — K. Alizadeh, S. I. Mirzadeh, et al., 2024 https://scholar.google.com/scholar?q=LLM+in+a+Flash%3A+Efficient+Large+Language+Model+Inference+with+Limited+Memory 14. PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving — A. C. Yüzügüler, J. Zhuang, L. Cavigelli, 2025 https://scholar.google.com/scholar?q=PRESERVE%3A+Prefetching+Model+Weights+and+KV-Cache+in+Distributed+LLM+Serving 15. MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs — H. Wu, Z. Cao, et al., 2026 https://scholar.google.com/scholar?q=MemExplorer%3A+Navigating+the+Heterogeneous+Memory+Design+Space+for+Agentic+Inference+NPUs 16. Lincoln: Real-Time 50~100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash Memory — W. Sun, M. Gao, et al., 2025 https://scholar.google.com/scholar?q=Lincoln%3A+Real-Time+50~100B+LLM+Inference+on+Consumer+Devices+with+LPDDR-Interfaced%2C+Compute-Enabled+Flash+Memory Interactive Visualization: FLINT: Sharing High Bandwidth Flash for Scalable LLM Inference
-
706
LinearKV: When Exact State Merging Breaks Hybrid Models
This episode examines LinearKV, a new approach to position-independent caching for hybrid Mamba-attention language models, and a striking result: the mathematically "correct" way to merge cached context — exact algebraic composition of recurrent states — badly degrades one tested model's output quality, while a simpler shortcut using only the most recent cached chunk performs reliably. The discussion contrasts standard prefix caching (used by vLLM's PagedAttention and SGLang's RadixAttention) with position-independent caching, which lets cached chunks be reused regardless of order, and explains why that trick breaks down for linear-recurrent layers like Mamba-2 and Gated DeltaNet, which compress history into a single fixed-size state rather than a token-indexed KV ledger. It traces the core problem to how each cached chunk's recurrent state was built in isolation, making "exact" composition exact only relative to a flawed reference rather than the true full-context computation. Listeners interested in LLM serving infrastructure, caching systems, or the tradeoffs of hybrid architectures will find this a concrete case study in how systems intuitions from full attention can actively mislead when applied to newer recurrent designs. Sources: 1. LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs — Yirui Liu, Ruoling Qi, Longwen Wang, Xuaner Wu, Jian Chen, Yuxin Jin, Jiawei Shao, Xuelong Li, 2026 http://arxiv.org/abs/2608.11231 2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, Ion Stoica, 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 3. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Ying Sheng, 2024 https://scholar.google.com/scholar?q=SGLang%3A+Efficient+Execution+of+Structured+Language+Model+Programs 4. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, Junchen Jiang, 2025 https://scholar.google.com/scholar?q=CacheBlend%3A+Fast+Large+Language+Model+Serving+for+RAG+with+Cached+Knowledge+Fusion 5. EPIC: Efficient Position-Independent Context Caching for Serving Large Language Models — Junhao Hu et al., 2025 https://scholar.google.com/scholar?q=EPIC%3A+Efficient+Position-Independent+Context+Caching+for+Serving+Large+Language+Models 6. HYPIC: Accelerating hybrid-attention LLM serving with position-independent caching — Yifei Liu, Juntong Wu, Yang Liu, Junhao Hu, Minghao Li, Xiaoxu Chen, Weihang Chen, 2026 https://scholar.google.com/scholar?q=HYPIC%3A+Accelerating+hybrid-attention+LLM+serving+with+position-independent+caching 7. Marconi: Prefix caching for the era of hybrid LLMs — Rui Pan, Zhuang Wang, Zhen Jia, Can Karakus, Luca Zancato, Tri Dao, Yida Wang, Ravi Netravali, 2025 https://scholar.google.com/scholar?q=Marconi%3A+Prefix+caching+for+the+era+of+hybrid+LLMs 8. Gated Delta Networks: Improving Mamba2 with the Delta Rule — Songlin Yang, Jan Kautz, Ali Hatamizadeh, 2025 (ICLR) https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+the+Delta+Rule 9. EPIC: Efficient position-independent caching for serving large language models — Junhao Hu, Wenrui Huang, Weidong Wang, et al., 2025 (ICML) https://scholar.google.com/scholar?q=EPIC%3A+Efficient+position-independent+caching+for+serving+large+language+models Interactive Visualization: LinearKV: When Exact State Merging Breaks Hybrid Models
-
705
Decoupling KL Direction from Rollout Source in LLM Distillation
This episode examines "Decoupling KL and Trajectories," which challenges the assumption that off-policy training must pair with forward KL divergence and on-policy training with reverse KL — a coupling used by DeepSeek-R1, Qwen3, MiMo, and GLM-5 without ever being tested. The hosts unpack the two independent design axes at play: prefix source (teacher-generated versus student-generated rollouts) and KL direction (forward's distribution-covering behavior versus reverse's mode-seeking concentration), showing that standard SFT is actually just forward-KL distillation on teacher text in disguise. The paper's real contribution is exploring two previously unstudied combinations — teacher-prefix reverse KL and student-prefix forward KL — turning an assumed two-option choice into a full four-quadrant design space. Grounding the discussion is a concrete experimental setup using Qwen3-0.6B-Base as a student distilling from Qwen3-4B and Qwen3-8B teachers. Listeners interested in LLM distillation, on-policy training, or exposure bias will find this a sharp dismantling of an industry-wide convention nobody had thought to question. Sources: 1. Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation — Anhao Zhao, Haoran Xin, Yingqi Fan, Junlong Tong, Wenjie Li, Xiaoyu Shen, 2026 http://arxiv.org/abs/2605.16826 2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning 3. Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks — Samy Bengio, Oriol Vinyals, Navdeep Jaitly, Noam Shazeer, 2015 https://scholar.google.com/scholar?q=Scheduled+Sampling+for+Sequence+Prediction+with+Recurrent+Neural+Networks 4. Sequence Level Training with Recurrent Neural Networks — Marc'Aurelio Ranzato, Sumit Chopra, Michael Auli, Wojciech Zaremba, 2016 https://scholar.google.com/scholar?q=Sequence+Level+Training+with+Recurrent+Neural+Networks 5. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, Olivier Bachem (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes 6. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2024 https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models 7. Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models — Feng Luo, Yu-Neng Chuang, Guanchu Wang, Zicheng Xu, Xiaotian Han, Tianyi Zhang, Vladimir Braverman, 2026 https://scholar.google.com/scholar?q=Demystifying+OPD%3A+Length+Inflation+and+Stabilization+Strategies+for+Large+Language+Models 8. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey J. Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29 9. Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting — Xinyuan Chen, Noam Razin, Karthik Narasimhan, Danqi Chen, 2025 https://scholar.google.com/scholar?q=Retaining+by+Doing%3A+The+Role+of+On-Policy+Data+in+Mitigating+Forgetting Interactive Visualization: Decoupling KL Direction from Rollout Source in LLM Distillation
-
704
Token Teachability: Rethinking Disagreement in On-Policy Distillation
This episode examines "Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation," which challenges a core assumption in on-policy knowledge distillation: that raw KL divergence between teacher and student token predictions is a reliable signal for which tokens deserve training focus. The discussion traces the lineage from Hinton's original distillation work through Google DeepMind's on-policy approach, then explains the paper's key insight — large disagreement can mean either a small, actionable correction the student can use, or a "incompatible" mismatch pointing toward options the student assigns near-zero probability, and raw KL can't distinguish the two. Building on this distinction, the authors introduce "token teachability" as a better selection criterion and a method called TA-OPD that trains only on the most teachable tokens, reportedly matching or beating full-dataset training while using just 5% of the tokens. Listeners interested in efficient model training, distillation techniques, or the gap between statistical salience and actual learnability will find the reframing of a decade-old assumption particularly compelling. Sources: 1. Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation — Yuanyi Wang, Su Lu, Yanggan Gu, Pengkai Wang, Yifan Yang, Zhaoyi Yan, Congkai Xie, Jianmin Wu, Hongxia Yang, 2026 http://arxiv.org/abs/2605.26844 2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015 https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network 3. Sequence-Level Knowledge Distillation — Yoon Kim, Alexander M. Rush, 2016 https://scholar.google.com/scholar?q=Sequence-Level+Knowledge+Distillation 4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, et al. (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes 5. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2024 https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models 6. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey J. Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29 7. Not All Tokens Are What You Need for Pretraining (Rho-1) — Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu, Yelong Shen, Ruochen Xu, Chen Lin, Yujiu Yang, Jian Jiao, Nan Duan, Weizhu Chen, 2024 https://scholar.google.com/scholar?q=Not+All+Tokens+Are+What+You+Need+for+Pretraining+%28Rho-1%29 8. Contrastive Decoding: Open-ended Text Generation as Optimization — Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, Mike Lewis, 2023 https://scholar.google.com/scholar?q=Contrastive+Decoding%3A+Open-ended+Text+Generation+as+Optimization 9. DistiLLM: Towards Streamlined Distillation for Large Language Models — Jongwoo Ko, Sungnyun Kim, Tianyi Chen, Se-Young Yun, 2024 https://scholar.google.com/scholar?q=DistiLLM%3A+Towards+Streamlined+Distillation+for+Large+Language+Models 10. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023 https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding 11. TIP: Token Importance in On-Policy Distillation — Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang, Alborz Geramifard, 2026 https://scholar.google.com/scholar?q=TIP%3A+Token+Importance+in+On-Policy+Distillation 12. Entropy-Aware On-Policy Distillation of Language Models — Woogyeol Jin, Taywon Min, Yongjin Yang, Swanand Ravindra Kadhe, Yi Zhou, Dennis Wei, Nathalie Baracaldo, Kimin Lee, 2026 https://scholar.google.com/scholar?q=Entropy-Aware+On-Policy+Distillation+of+Language+Models 13. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe — Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, et al., 2026 https://scholar.google.com/scholar?q=Rethinking+On-Policy+Distillation+of+Large+Language+Models%3A+Phenomenology%2C+Mechanism%2C+and+Recipe 14. Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning — Shenzhi Wang, Le Yu, Chang Gao, Chujie Zheng, Shixuan Liu, Rui Lu, et al., 2026 https://scholar.google.com/scholar?q=Beyond+the+80%2F20+Rule%3A+High-Entropy+Minority+Tokens+Drive+Effective+Reinforcement+Learning+for+LLM+Reasoning 15. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — Daya Guo, Dejian Yang, Haowei Zhang, et al., 2025 https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning Interactive Visualization: Token Teachability: Rethinking Disagreement in On-Policy Distillation
-
703
Weak-to-Strong On-Policy Distillation Beats the Teacher
This episode examines Weak-to-Strong On-Policy Distillation, a paper from University of Maryland, Microsoft Research, and MBZUAI showing that an 8-billion-parameter student model can outperform every teacher used to train it on math benchmarks. The discussion traces the method's lineage from Hinton's original 2015 knowledge distillation through DAgger's 2011 on-policy correction idea to 2023 weak-to-strong generalization work from OpenAI's Superalignment team, explaining how the paper inverts DAgger's core assumption by using a supervisor weaker than the student rather than a trusted expert. It also contrasts this approach with reinforcement learning from verifiable rewards, which gives only a sparse end-of-rollout signal, versus on-policy distillation's dense per-token feedback on the student's own generated trajectories. The hosts situate the work against industry precedents like Qwen3's large-to-small distillation and multi-teacher on-policy distillation, both of which still require a teacher at least as capable as the student — a constraint this paper's method aims to eliminate. Listeners interested in how weaker models might keep training stronger ones as the field runs out of better supervisors will find the mechanism and its research lineage laid out in detail. Sources: 1. Weak-to-Strong On-Policy Distillation — Fangxu Yu, Weijia Xu, Michael Xu, Tianyi Zhou, Zinan Lin, 2026 http://arxiv.org/abs/2607.26246 2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015 https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network 3. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29 4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (Generalized Knowledge Distillation, GKD) — Rishabh Agarwal, Nino Vieillard, et al. (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes+%28Generalized+Knowledge+Distillation%2C+GKD%29 5. On-Policy Distillation (blog post / technical report) — Thinking Machines Lab (cited in the paper as 'Lu & Lab'), 2025 https://scholar.google.com/scholar?q=On-Policy+Distillation+%28blog+post+%2F+technical+report%29 6. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision — Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, Jeff Wu, 2023 https://scholar.google.com/scholar?q=Weak-to-Strong+Generalization%3A+Eliciting+Strong+Capabilities+With+Weak+Supervision 7. Weak-to-Strong Generalization beyond Accuracy: a Roadmap in Codegen, Safety, and Beyond (or closely related 2024 weak-to-strong follow-up work, cited in the paper as 'Ning et al., 2024') — Ning et al., 2024 https://scholar.google.com/scholar?q=Weak-to-Strong+Generalization+beyond+Accuracy%3A+a+Roadmap+in+Codegen%2C+Safety%2C+and+Beyond+%28or+closely+related+2024+weak-to-strong+follow-up+work%2C+cited+in+the+paper+as+%27Ning+et+al.%2C+2024%27%29 8. Illustrating Reinforcement Learning from Human Feedback / scalable oversight lineage (e.g. Christiano et al., 'Deep Reinforcement Learning from Human Preferences', and related scalable-oversight work) — Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, Dario Amodei (RLHF); related scalable oversight literature, 2017 (RLHF) / ongoing scalable oversight literature https://scholar.google.com/scholar?q=Illustrating+Reinforcement+Learning+from+Human+Feedback+%2F+scalable+oversight+lineage+%28e.g.+Christiano+et+al.%2C+%27Deep+Reinforcement+Learning+from+Human+Preferences%27%2C+and+related+scalable-oversight+work%29 9. Contrastive Decoding: Open-ended Text Generation as Optimization — Li, Holtzman, Fried, Liang, Eisner, Hashimoto, Zettlemoyer, Lewis, 2023 https://scholar.google.com/scholar?q=Contrastive+Decoding%3A+Open-ended+Text+Generation+as+Optimization 10. Distillation Scaling Laws — Busbridge, Ramapuram, Ablin, et al. (Apple), 2024 https://scholar.google.com/scholar?q=Distillation+Scaling+Laws 11. Speculative Decoding papers (e.g., Leviathan et al., 'Fast Inference from Transformers via Speculative Decoding', 2023) — Leviathan, Kalman, Matias, 2023 https://scholar.google.com/scholar?q=Speculative+Decoding+papers+%28e.g.%2C+Leviathan+et+al.%2C+%27Fast+Inference+from+Transformers+via+Speculative+Decoding%27%2C+2023%29 Interactive Visualization: Weak-to-Strong On-Policy Distillation Beats the Teacher
-
702
On-Policy Distillation: Why a Stronger Teacher Can Backfire
This episode examines why on-policy distillation (OPD) of large language models can fail catastrophically even when the teacher model is objectively stronger by every benchmark — a 7-billion-parameter teacher completely failed to improve a 1.5-billion-parameter student, while a smaller, weaker teacher succeeded. The discussion traces OPD's mechanics: unlike classic distillation, which trains students on teacher-generated text and suffers from exposure bias, OPD has the student generate its own rollouts and scores them against the teacher's token-level probability distribution via reverse KL divergence, yielding a dense per-token reward without a verifier. The central finding is that teacher-student "overlap ratio" — how much the teacher's likely next tokens actually match the student's own candidate set — determines whether that dense signal teaches anything, meaning benchmark strength and teachability are fundamentally different axes. The paper's three-part structure (phenomenology, mechanism, recipe) is illustrated through a controlled comparison of two similarly-scored Qwen3-4B teacher variants, isolating thinking-pattern compatibility as the real driver of distillation success. Listeners working on post-training or distillation pipelines will find this a direct challenge to the common assumption that upgrading the teacher model is a safe, unconditional improvement. Sources: 1. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe — Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, Tianyu Yu, Huan-ang Gao, Wenkai Yang, Zhiyuan Liu, Ning Ding, 2026 http://arxiv.org/abs/2604.13016 2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning 3. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2023 https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models 4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (GKD) — Rishabh Agarwal, Nino Vieillard, and colleagues (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes+%28GKD%29 5. Qwen3 Technical Report — An Yang and the Qwen Team (Alibaba), 2025 https://scholar.google.com/scholar?q=Qwen3+Technical+Report 6. On-policy distillation of language models: Learning from self-generated mistakes (MiniLLM) — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2023 https://scholar.google.com/scholar?q=On-policy+distillation+of+language+models%3A+Learning+from+self-generated+mistakes+%28MiniLLM%29 7. Distillation scaling laws — Dan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram, Etai Littwin, Russ Webb, 2025 https://scholar.google.com/scholar?q=Distillation+scaling+laws 8. On the efficacy of knowledge distillation — Jang Hyun Cho, Bharath Hariharan, 2019 https://scholar.google.com/scholar?q=On+the+efficacy+of+knowledge+distillation 9. Small models struggle to learn from strong reasoners — Yuetai Li, Xiang Yue, Zhangchen Xu, Fengqing Jiang, Luyao Niu, Bill Yuchen Lin, Bhaskar Ramasubramanian, Radha Poovendran, 2025 https://scholar.google.com/scholar?q=Small+models+struggle+to+learn+from+strong+reasoners 10. On-policy distillation (Thinking Machines Lab blog) — Kevin Lu and Thinking Machines Lab, 2025 https://scholar.google.com/scholar?q=On-policy+distillation+%28Thinking+Machines+Lab+blog%29 11. Learning beyond teacher: Generalized on-policy distillation with reward extrapolation — Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang, Saiyong Yang, Yankai Lin, 2026 https://scholar.google.com/scholar?q=Learning+beyond+teacher%3A+Generalized+on-policy+distillation+with+reward+extrapolation Interactive Visualization: On-Policy Distillation: Why a Stronger Teacher Can Backfire
-
701
Sleeper Memory Poisoning: When Assistants Remember Lies
This episode examines "Hidden in Memory: Sleeper Memory Poisoning in LLM Agents," a study showing how attackers can plant fabricated facts into an AI assistant's persistent memory that lie dormant until triggered in an unrelated future conversation. Unlike traditional prompt injection, which manipulates a model's behavior only within a single session, this attack targets the memory-write step itself, allowing a single black-box, universal payload template—refined through an actor-critic search between attacker and critic LLMs—to succeed across arbitrary goals with startlingly high rates (up to 99.8% on GPT-5.5). The discussion breaks down the three-stage pipeline attackers must clear (injection, retrieval, and usage), and highlights a clever technique for maximizing the odds a poisoned memory resurfaces later: rewriting it to boost embedding similarity with plausible future queries while a semantic-consistency judge guards against the rewrite drifting from the original intent. Testing spans 700 document-goal pairs across 15 source types and multiple commercial memory architectures, revealing that whether the model or a separate manager process controls memory writes dramatically changes how exploitable a system is. It's a sobering look at how "memory" — now a default feature across ChatGPT, Claude, Gemini, and agent frameworks like Mem0 — introduces a persistent, hard-to-detect attack surface that outlives the malicious content that created it. Sources: 1. Hidden in Memory: Sleeper Memory Poisoning in LLM Agents — Sidharth Pulipaka, Stanislau Hlebik, Leonidas Raghav, Sahar Abdelnabi, Vyas Raina, Ivaxi Sheth, Mario Fritz, 2026 http://arxiv.org/abs/2605.15338 2. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training — Evan Hubinger et al. (Anthropic), 2024 https://scholar.google.com/scholar?q=Sleeper+Agents%3A+Training+Deceptive+LLMs+that+Persist+Through+Safety+Training 3. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection — Kai Greshake, Sahar Abdelnabi, et al., 2023 https://scholar.google.com/scholar?q=Not+What+You%27ve+Signed+Up+For%3A+Compromising+Real-World+LLM-Integrated+Applications+with+Indirect+Prompt+Injection 4. Prompt Injection Attack Against LLM-Integrated Applications — Yi Liu et al., 2023 https://scholar.google.com/scholar?q=Prompt+Injection+Attack+Against+LLM-Integrated+Applications 5. Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory — Prateek Chhikara et al., 2025 https://scholar.google.com/scholar?q=Mem0%3A+Building+Production-Ready+AI+Agents+with+Scalable+Long-Term+Memory 6. Universal and Transferable Adversarial Attacks on Aligned Language Models (GCG) — Zou et al., 2023 https://scholar.google.com/scholar?q=Universal+and+Transferable+Adversarial+Attacks+on+Aligned+Language+Models+%28GCG%29 7. AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases — Chen, Xiang, Xiao, Song, Li, 2024 (NeurIPS) https://scholar.google.com/scholar?q=AgentPoison%3A+Red-teaming+LLM+Agents+via+Poisoning+Memory+or+Knowledge+Bases 8. GEPA: Efficient Textual Optimization via LLM-based Reflection and Pareto-Efficient Evolutionary Search — Agrawal, Khattab, Potts, 2025 https://scholar.google.com/scholar?q=GEPA%3A+Efficient+Textual+Optimization+via+LLM-based+Reflection+and+Pareto-Efficient+Evolutionary+Search 9. Injection through web agents that fetch pages with hidden HTML instructions — Raghav and Choong, 2026 https://scholar.google.com/scholar?q=Injection+through+web+agents+that+fetch+pages+with+hidden+HTML+instructions 10. The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers — Bullwinkel, Severi, Hines, Minnich, Kumar, Zunger, 2026 https://scholar.google.com/scholar?q=The+Trigger+in+the+Haystack%3A+Extracting+and+Reconstructing+LLM+Backdoor+Triggers Interactive Visualization: Sleeper Memory Poisoning: When Assistants Remember Lies
-
700
Why AI Systems Don't Learn After Deployment
This episode examines a paper by Emmanuel Dupoux, Yann LeCun, and Jitendra Malik arguing that deployed AI models learn nothing after training, unlike a toddler who continuously experiments through action, observation, imitation, and inquiry. The discussion breaks down the paper's core distinction between System A (passive, observation-based statistical learning like self-supervised training) and System B (action-based reinforcement learning through feedback), and explains why neither alone can produce autonomous intelligence. It then covers the paper's proposed fix, System M, an orchestrator modeled on software-defined networking that monitors low-bandwidth "meta-state" signals like prediction error and confidence to dynamically route between learning systems, automating what human MLOps engineers currently do by hand. The conversation also connects this framework to LeCun's 2022 autonomous machine intelligence proposal and the ongoing debate sparked by Silver and Sutton's "Era of Experience" critique about AI hitting a data wall. Listeners interested in the architecture of autonomous learning and what's actually missing between today's static models and genuinely adaptive intelligence will find the systems-level framing illuminating. Sources: 1. Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science — Emmanuel Dupoux, Yann LeCun, Jitendra Malik, 2026 http://arxiv.org/abs/2603.15381 2. A Path Towards Autonomous Machine Intelligence — Yann LeCun, 2022 https://scholar.google.com/scholar?q=A+Path+Towards+Autonomous+Machine+Intelligence 3. Welcome to the Era of Experience — David Silver, Richard Sutton, 2025 https://scholar.google.com/scholar?q=Welcome+to+the+Era+of+Experience 4. Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (MuZero) — Julian Schrittwieser et al., 2020 https://scholar.google.com/scholar?q=Mastering+Atari%2C+Go%2C+Chess+and+Shogi+by+Planning+with+a+Learned+Model+%28MuZero%29 5. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning — Mido Assran et al. (incl. LeCun), 2025 https://scholar.google.com/scholar?q=V-JEPA+2%3A+Self-Supervised+Video+Models+Enable+Understanding%2C+Prediction+and+Planning 6. Coordination Among Neural Modules Through a Shared Global Workspace — Anirudh Goyal, Aniket Didolkar, et al., 2022 https://scholar.google.com/scholar?q=Coordination+Among+Neural+Modules+Through+a+Shared+Global+Workspace 7. Embodied AI Agents: Modeling the World — Pascale Fung, Emmanuel Dupoux, Jitendra Malik, et al., 2025 https://scholar.google.com/scholar?q=Embodied+AI+Agents%3A+Modeling+the+World Interactive Visualization: Why AI Systems Don't Learn After Deployment
-
699
Second-Order Optimization Meets Runtime Scheduling at Scale
This episode examines "Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training," a May 2026 Oxford paper introducing Asteria, a runtime system rather than a new optimizer. The discussion covers why curvature-aware methods like Shampoo and SOAP have never displaced AdamW despite converging in fewer steps, tracing the lineage from K-FAC through Distributed Shampoo to SOAP and explaining the Kronecker-factorization tricks that make tracking curvature tractable at all. The hosts unpack the paper's "three physical walls" framework — a vertical capacity wall from single-GPU memory limits, an overlap disruption wall where cubic-cost matrix operations stall compute-communication overlap, and a global consensus wall from synchronous full-state updates across mismatched network speeds — and debate whether reengineering the plumbing around an unchanged optimizer counts as a genuine research contribution. Listeners interested in distributed training infrastructure, optimizer design trade-offs, or the gap between algorithmic elegance and practical deployability will find the back-and-forth over real benchmark numbers (96 seconds versus 1.5 seconds per step) especially grounded. Sources: 1. Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training — Yishun Lu, Junhao Zhang, Zeyu Yang, Wes Armour, 2026 http://arxiv.org/abs/2605.16184 2. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature 3. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization 4. Scalable Second Order Optimization for Deep Learning — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer, 2020 https://scholar.google.com/scholar?q=Scalable+Second+Order+Optimization+for+Deep+Learning 5. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira, David Brandfonbrener, Sham Kakade, et al., 2024 https://scholar.google.com/scholar?q=SOAP%3A+Improving+and+Stabilizing+Shampoo+using+Adam 6. A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale — H.-J. M. Shi, T.-H. Lee, S. Iwasaki, J. Gallego-Posada, Z. Li, K. Rangadurai, D. Mudigere, M. Rabbat, 2023 https://scholar.google.com/scholar?q=A+Distributed+Data-Parallel+PyTorch+Implementation+of+the+Distributed+Shampoo+Optimizer+for+Training+Neural+Networks+At-Scale 7. Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading — A. Maurya, J. Ye, M. M. Rafique, F. Cappello, B. Nicolae, 2024 https://scholar.google.com/scholar?q=Deep+Optimizer+States%3A+Towards+Scalable+Training+of+Transformer+Models+Using+Interleaved+Offloading 8. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith, Y. He, 2021 https://scholar.google.com/scholar?q=ZeRO-Infinity%3A+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning 9. Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization — W. Lin, S. C. Lowe, F. Dangel, R. Eschenhagen, Z. Xu, R. B. Grosse, 2026 https://scholar.google.com/scholar?q=Understanding+and+Improving+Shampoo+and+SOAP+via+Kullback-Leibler+Minimization 10. Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training — Y. Lu, W. Armour, 2026 https://scholar.google.com/scholar?q=Beyond+the+Mean%3A+Fisher-Orthogonal+Projection+for+Natural+Gradient+Descent+in+Large+Batch+Training Interactive Visualization: Second-Order Optimization Meets Runtime Scheduling at Scale
-
698
Hand-Written PTX vs WMMA: A Precision-Dependent GPU Speedup
This episode digs into a paper testing whether hand-written PTX assembly beats NVIDIA's WMMA API for Tensor Core GEMM kernels on an L4 GPU, finding that the answer flips depending on numeric precision rather than holding as a universal rule. The discussion covers the hardware distinction between Tensor Cores and regular CUDA cores, and contrasts the convenience of WMMA against the finer control PTX offers through instructions like cp.async, ldmatrix, and mma.sync. A key thread traces why this matters in practice: quantized LLM serving at INT8 or INT4 shifts kernels from compute-bound to memory-bound, making the precision-dependent payoff of hand-tuned PTX directly relevant to running open-weight models cheaply. The episode also addresses the methodological choice to test on a single GPU, arguing that isolating precision and instruction-set effects requires holding hardware constant rather than spreading across devices. Listeners get a concrete framework for deciding when the extra engineering effort of writing raw PTX is worth it versus when it's wasted work. Sources: 1. Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4 — Matt J. Borowski, Blazej Osinski, 2026 http://arxiv.org/abs/2608.10103 2. NVIDIA Tensor Core Programmability, Performance & Precision — Stefano Markidis, Steven W. D. Chien, Erwin Laure, Ivy B. Peng, Jeffrey S. Vetter, 2018 https://scholar.google.com/scholar?q=NVIDIA+Tensor+Core+Programmability%2C+Performance+%26+Precision 3. Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking — Zhe Jia, Marco Maggioni, Jeffrey Smith, Daniele Paolo Scarpazza, 2018 https://scholar.google.com/scholar?q=Dissecting+the+NVIDIA+Volta+GPU+Architecture+via+Microbenchmarking 4. CUTLASS: CUDA Templates for Linear Algebra Subroutines — Andrew Kerr, Duane Merrill, Julien Demouth, John Tran (NVIDIA), with ongoing project contributors, 2018 (initial release, actively maintained since) https://scholar.google.com/scholar?q=CUTLASS%3A+CUDA+Templates+for+Linear+Algebra+Subroutines 5. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré, 2022 https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness 6. Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations — Philippe Tillet, H. T. Kung, David Cox, 2019 https://scholar.google.com/scholar?q=Triton%3A+An+Intermediate+Language+and+Compiler+for+Tiled+Neural+Network+Computations 7. Ansor: Generating High-Performance Tensor Programs for Deep Learning — Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, Ion Stoica, 2020 (OSDI) https://scholar.google.com/scholar?q=Ansor%3A+Generating+High-Performance+Tensor+Programs+for+Deep+Learning 8. Dissecting the Ampere GPU Architecture via Microbenchmarking — Wei Sun, Ang Li, Tong Geng, Sander Stuijk, Henk Corporaal, 2022 https://scholar.google.com/scholar?q=Dissecting+the+Ampere+GPU+Architecture+via+Microbenchmarking 9. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning — Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy, 2018 (OSDI) https://scholar.google.com/scholar?q=TVM%3A+An+Automated+End-to-End+Optimizing+Compiler+for+Deep+Learning 10. Understanding Latency Hiding on GPUs — V. Volkov, 2016 https://scholar.google.com/scholar?q=Understanding+Latency+Hiding+on+GPUs 11. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — J. Lin et al., 2024 https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration 12. FlashInfer: Kernel Library for LLM Serving — Z. Ye et al., 2024 https://scholar.google.com/scholar?q=FlashInfer%3A+Kernel+Library+for+LLM+Serving
-
697
Mechanist: Automating the Discovery of How AI Models Think
This episode examines "Mechanist," a multi-agent system built by researchers at Zhejiang University, NUS, Southern University of Science and Technology, Heriot-Watt, UC San Diego, and Northeastern University to automate mechanistic interpretability research itself, rather than automating experiments in an external domain like chemistry or biology. The discussion covers how a central orchestrator coordinates four agents—hypothesis, experiment, verification, and iteration—drawing on a 13,000-paper interpretability knowledge graph and a 43-million-paper cross-disciplinary graph called SciAtlas to generate and test theories about how models actually compute. Key concepts explored include subliminal learning, where a trait transfers from teacher to student model through data that looks unrelated to it, and a three-frame belief decomposition (World Knowledge, Personal Belief, Attributed Belief) used to probe whether models genuinely separate fact from attributed belief. The episode previews four escalating case studies, starting with the discovery of a previously unflagged multimodal safety risk and building toward using mechanistic theories to directly intervene on model internals and even steer a biological system, raising the question of whether an AI system can meaningfully explain the black box that produced it. Sources: 1. Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence — Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen, 2026 http://arxiv.org/abs/2608.12036 2. Language models transmit behavioural traits through hidden signals in data — Alex Cloud, Minh Le, James Chua, Jan Betley, Anna Sztyber-Betley, Sören Mindermann, Jacob Hilton, Samuel Marks, Owain Evans, 2026 (Nature) https://scholar.google.com/scholar?q=Language+models+transmit+behavioural+traits+through+hidden+signals+in+data 3. Subliminal learning is a lora artifact — Todd Nief, Harvey Yiyun Fu, Mark Muchane, Ari Holtzman, 2026 https://scholar.google.com/scholar?q=Subliminal+learning+is+a+lora+artifact 4. Language models cannot reliably distinguish belief from knowledge and fact — Mirac Suzgun, Tayfun Gur, Federico Bianchi, Daniel E. Ho, Thomas Icard, Dan Jurafsky, James Zou, 2025 (Nature Machine Intelligence) https://scholar.google.com/scholar?q=Language+models+cannot+reliably+distinguish+belief+from+knowledge+and+fact 5. Genome modelling and design across all domains of life with evo 2 — Garyk Brixi, Matthew G. Durrant, Jerome Ku, et al., 2026 (Nature) https://scholar.google.com/scholar?q=Genome+modelling+and+design+across+all+domains+of+life+with+evo+2 6. Towards end-to-end automation of ai research — Chris Lu, Cong Lu, Robert Tjarko Lange, Yutaro Yamada, Shengran Hu, Jakob Foerster, David Ha, Jeff Clune, 2026 (Nature) https://scholar.google.com/scholar?q=Towards+end-to-end+automation+of+ai+research 7. Sleeper agents: Training deceptive llms that persist through safety training — Evan Hubinger et al., 2024 https://scholar.google.com/scholar?q=Sleeper+agents%3A+Training+deceptive+llms+that+persist+through+safety+training 8. Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers — Jan Dubinski, Jan Betley, Anna Sztyber-Betley, Daniel Tan, Owain Evans, 2026 https://scholar.google.com/scholar?q=Conditional+misalignment%3A+common+interventions+can+hide+emergent+misalignment+behind+contextual+triggers
-
696
Inside Claude Code's Agentic Loop: A Design Space Analysis
This episode dissects the internal architecture of Claude Code by examining its extracted TypeScript source (v2.1.88) alongside two other agent systems, OpenClaw and Hermes Agent, drawing on the paper "Dive into Claude Code" by Jiacheng Liu et al. from VILA Lab at MBZUAI. It reveals that the model's reasoning core is essentially a single while-loop — literally called queryLoop() — with everything else (permissions, context management, tools, subagents) built as scaffolding around it. The discussion covers deny-first permission rules, the five-stage compaction pipeline for managing context windows, subagent delegation with isolated context windows, and how the Model Context Protocol connects to external tool servers. The hosts also extract five human values embedded directly in the code — human decision authority, safety/security/privacy, reliable execution, capability amplification, and contextual adaptability — framing the episode as less about how Claude Code works and more about what its designers chose to prioritize, made visible through actual implementation choices. Sources: 1. Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems — Jiacheng Liu, Xiaohan Zhao, Xinyi Shang, Zhiqiang Shen, 2026 http://arxiv.org/abs/2604.14228 2. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection — Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, Mario Fritz, 2023 https://scholar.google.com/scholar?q=Not+What+You%27ve+Signed+Up+For%3A+Compromising+Real-World+LLM-Integrated+Applications+with+Indirect+Prompt+Injection 3. GAIA: A Benchmark for General AI Assistants — Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, Thomas Scialom, 2023 https://scholar.google.com/scholar?q=GAIA%3A+A+Benchmark+for+General+AI+Assistants 4. Constitutional AI: Harmlessness from AI Feedback — Yuntao Bai, Saurav Kadavath, Sandipan Kundu, et al., 2022 https://scholar.google.com/scholar?q=Constitutional+AI%3A+Harmlessness+from+AI+Feedback 5. AgentBench: Evaluating LLMs as Agents — Xiao Liu, Hao Yu, Hanchen Zhang, et al., 2023 https://scholar.google.com/scholar?q=AgentBench%3A+Evaluating+LLMs+as+Agents
-
695
Grep vs Vector Search: How Agent Harnesses Shape Retrieval Accuracy
This episode examines "Is Grep All You Need? How Agent Harnesses Reshape Agentic Search," which tests whether simple regex-based retrieval can outperform vector search inside agentic pipelines like Chronos, Claude Code, and Codex CLI. The hosts dig into how retrieval mode interacts with harness architecture, delivery method (inline vs. programmatic), and backbone model choice, finding that inline grep beats inline vector search across every harness-model pairing tested — with gaps as wide as twenty points and swings as large as switching harnesses entirely. A striking case shows the same model scoring 93.1% on one harness but only 76.7% on another, suggesting orchestration and prompt construction matter as much as the retrieval algorithm itself. The discussion also surfaces a counterintuitive twist: forcing an agent to read retrieved results from a file instead of getting them dumped inline can nearly halve accuracy, even with identical underlying search. Listeners interested in RAG, agent design, or LLM evaluation methodology will find the paper's tangled-but-honest approach to measuring real deployed systems a useful corrective to cleaner but less realistic ablation studies. Sources: 1. Is Grep All You Need? How Agent Harnesses Reshape Agentic Search — Sahil Sen, Akhil Kasturi, Elias Lumer, Anmol Gulati, Vamse Kumar Subbiah, 2026 http://arxiv.org/abs/2605.15184 2. ReAct: Synergizing Reasoning and Acting in Language Models — Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao, 2022 https://scholar.google.com/scholar?q=ReAct%3A+Synergizing+Reasoning+and+Acting+in+Language+Models 3. WebGPT: Browser-assisted question-answering with human feedback — Reiichiro Nakano, Jacob Hilton, Suchir Balaji, et al. (OpenAI), 2021 https://scholar.google.com/scholar?q=WebGPT%3A+Browser-assisted+question-answering+with+human+feedback 4. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Patrick Lewis, Ethan Perez, Aleksandra Piktus, et al. (Facebook AI Research), 2020 https://scholar.google.com/scholar?q=Retrieval-Augmented+Generation+for+Knowledge-Intensive+NLP+Tasks 5. LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory — Di Wu, Hongwei Wang, Wenhao Yu, et al., 2024 https://scholar.google.com/scholar?q=LongMemEval%3A+Benchmarking+Chat+Assistants+on+Long-Term+Interactive+Memory 6. The Probabilistic Relevance Framework: BM25 and Beyond — Stephen Robertson, Hugo Zaragoza, 2009 https://scholar.google.com/scholar?q=The+Probabilistic+Relevance+Framework%3A+BM25+and+Beyond 7. Dense Passage Retrieval for Open-Domain Question Answering — Vladimir Karpukhin, Barlas Oğuz, Sewon Min, et al. (Facebook AI Research), 2020 https://scholar.google.com/scholar?q=Dense+Passage+Retrieval+for+Open-Domain+Question+Answering 8. Lost in the Middle: How Language Models Use Long Contexts — Nelson F. Liu, Kevin Lin, John Hewitt, et al., 2023 https://scholar.google.com/scholar?q=Lost+in+the+Middle%3A+How+Language+Models+Use+Long+Contexts 9. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking — Thibault Formal, Benjamin Piwowarski, Stéphane Clinchant, 2021 https://scholar.google.com/scholar?q=SPLADE%3A+Sparse+Lexical+and+Expansion+Model+for+First+Stage+Ranking 10. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models — Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, Iryna Gurevych, 2021 https://scholar.google.com/scholar?q=BEIR%3A+A+Heterogenous+Benchmark+for+Zero-shot+Evaluation+of+Information+Retrieval+Models 11. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Vivian Fang, Shishir G. Patil, Kevin Lin, Sarah Wooders, Joseph E. Gonzalez, 2023 https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems 12. Chronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term Memory — Sahil Sen, Elias Lumer, Anmol Gulati, Vamse Kumar Subbiah, 2026 https://scholar.google.com/scholar?q=Chronos%3A+Temporal-Aware+Conversational+Agents+with+Structured+Event+Retrieval+for+Long-Term+Memory Interactive Visualization: Grep vs Vector Search: How Agent Harnesses Shape Retrieval Accuracy
-
694
Decomposing Speedups Across Runtime, Kernel, and Quantization
This episode dissects a common but misleading claim in LLM serving benchmarks: that swapping to a faster stack alone explains a headline speedup number. It walks through the mechanics separating prefill's compute-bound math from decode's memory-bandwidth-bound token generation, then explains how continuous batching and PagedAttention keep GPUs saturated, and how GPTQ quantization paired with the Marlin kernel avoids costly dequantization overhead. By introducing a matched FP16 intermediate stack, the paper cleanly splits a reported speedup into a runtime factor and a kernel-plus-quantization factor that multiply back to the observed total. The headline finding is striking: a controlled, apples-to-apples comparison yields a 2.58x speedup dominated by runtime (68%), while a more common operational comparison — pitting a batch-capped baseline against a fully loaded modern stack — inflates that to 10.62x, with runtime's share climbing to 86%. Listeners get a rare, rigorous look at how much of the "free lunch" in inference speedups is real architecture versus benchmark framing. Sources: 1. Decomposing Speedups Across Runtime, Kernel, and Quantization https://arxiv.org/pdf/2607.11368 2. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, et al., 2023 https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU 3. DeepSpeed-Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale — Reza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, et al., 2022 https://scholar.google.com/scholar?q=DeepSpeed-Inference%3A+Enabling+Efficient+Inference+of+Transformer+Models+at+Unprecedented+Scale 4. SqueezeLLM: Dense-and-Sparse Quantization — Sehoon Kim, Coleman Hooper, Amir Gholami, et al., 2023 https://scholar.google.com/scholar?q=SqueezeLLM%3A+Dense-and-Sparse+Quantization 5. Blink: Fast and Generic Collectives for Distributed ML — Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee, et al., 2020 https://scholar.google.com/scholar?q=Blink%3A+Fast+and+Generic+Collectives+for+Distributed+ML Interactive Visualization: Decomposing Speedups Across Runtime, Kernel, and Quantization
-
693
AI-Generated Text and the Death of the Open Web
This episode examines a study analyzing the growth of AI-generated content across the open web, drawing on 33 monthly samples from the Internet Archive's Wayback Machine between August 2022 and May 2025. It highlights the paper's central finding that AI-generated or AI-assisted content on newly published websites rose from zero before ChatGPT's launch to roughly 35 percent by mid-2025, and explores how the authors transform "Dead Internet Theory" from internet folklore into six testable hypotheses — including semantic contraction, truth decay, positivity shift, epistemic islands, entropy dilution, and stylistic monoculture. The discussion covers the methodology behind sampling a representative slice of the internet, including logarithmic downsampling and stratification across time, MIME type, and domain to avoid bias toward heavily-crawled sites. It also connects the findings to the concept of model collapse, framing the 35 percent figure as empirical evidence for a previously theoretical concern about AI models training on their own synthetic output. Listeners interested in web ecosystem health, LLM training data quality, or the intersection of internet culture and rigorous data science will find the episode's blend of meme-to-metric translation particularly compelling. Sources: 1. The Impact of AI-Generated Text on the Internet — Jonas Dolezal, Sawood Alam, Mark Graham, Maty Bohacek, 2026 http://arxiv.org/abs/2604.26965 2. DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature — Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, Chelsea Finn, 2023 https://scholar.google.com/scholar?q=DetectGPT%3A+Zero-Shot+Machine-Generated+Text+Detection+using+Probability+Curvature 3. Can AI-Generated Text be Reliably Detected? — Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, Soheil Feizi, 2023 https://scholar.google.com/scholar?q=Can+AI-Generated+Text+be+Reliably+Detected%3F 4. A Watermark for Large Language Models — John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, Tom Goldstein, 2023 https://scholar.google.com/scholar?q=A+Watermark+for+Large+Language+Models 5. GLTR: Statistical Detection and Visualization of Generated Text — Sebastian Gehrmann, Hendrik Strobelt, Alexander M. Rush, 2019 https://scholar.google.com/scholar?q=GLTR%3A+Statistical+Detection+and+Visualization+of+Generated+Text 6. The Curse of Recursion: Training on Generated Data Makes Models Forget — Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, Ross Anderson, 2024 (Nature, updated from 2023 preprint) https://scholar.google.com/scholar?q=The+Curse+of+Recursion%3A+Training+on+Generated+Data+Makes+Models+Forget 7. Is the Internet Dead? Evaluating Claims about Web Homogenization from AI-Generated Content (representative of the broader 2024-2025 measurement literature, e.g. Muzumdar et al. and related environmental-scanning studies of Dead Internet Theory discourse) — Various (this cluster of work is cited in the paper as Muzumdar et al., 2025 and similar), 2025 https://scholar.google.com/scholar?q=Is+the+Internet+Dead%3F+Evaluating+Claims+about+Web+Homogenization+from+AI-Generated+Content+%28representative+of+the+broader+2024-2025+measurement+literature%2C+e.g.+Muzumdar+et+al.+and+related+environmental-scanning+studies+of+Dead+Internet+Theory+discourse%29 8. Studies on bot/inauthentic-content prevalence on specific platforms (e.g., La Cava et al. 2025 on social media, Matatov et al. 2024 on platform-specific AI content) — Lucio La Cava et al.; Jonathan Matatov et al. (representative platform-specific studies cited in the paper's related-work section), 2024-2025 https://scholar.google.com/scholar?q=Studies+on+bot%2Finauthentic-content+prevalence+on+specific+platforms+%28e.g.%2C+La+Cava+et+al.+2025+on+social+media%2C+Matatov+et+al.+2024+on+platform-specific+AI+content%29 9. Public Trust and Perceptions of Artificial Intelligence (Ipsos / Reuters Institute Digital News Report and Edelman Trust Barometer AI-focused editions) — Ipsos (various); Reuters Institute for the Study of Journalism (Nic Newman et al.); Edelman Trust Barometer team, 2023-2025 (recurring annual) https://scholar.google.com/scholar?q=Public+Trust+and+Perceptions+of+Artificial+Intelligence+%28Ipsos+%2F+Reuters+Institute+Digital+News+Report+and+Edelman+Trust+Barometer+AI-focused+editions%29 10. Americans' Views of Artificial Intelligence (Pew Research Center recurring survey series) — Pew Research Center (Alec Tyson, Emma Kikuchi, and colleagues), 2023-2025 (recurring) https://scholar.google.com/scholar?q=Americans%27+Views+of+Artificial+Intelligence+%28Pew+Research+Center+recurring+survey+series%29 11. The Perception Gap: Comparing Public Beliefs about Misinformation to Empirical Prevalence Estimates (representative of the risk-perception vs. measured-prevalence literature this paper's framing descends from, e.g. work following Duffy et al. and general misperception-of-misinformation-prevalence studies) — Andrew Guess, colleagues in the misinformation-prevalence research cluster (representative of this line, distinct from the AI-specific surveys above), 2019-2023 (foundational misinformation-perception literature) https://scholar.google.com/scholar?q=The+Perception+Gap%3A+Comparing+Public+Beliefs+about+Misinformation+to+Empirical+Prevalence+Estimates+%28representative+of+the+risk-perception+vs.+measured-prevalence+literature+this+paper%27s+framing+descends+from%2C+e.g.+work+following+Duffy+et+al.+and+general+misperception-of-misinformation-prevalence+studies%29 12. Documenting the English Colossal Clean Crawled Corpus (C4) — Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner, 2021 https://scholar.google.com/scholar?q=Documenting+the+English+Colossal+Clean+Crawled+Corpus+%28C4%29 13. The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only — Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, Julien Launay, 2023 https://scholar.google.com/scholar?q=The+RefinedWeb+Dataset+for+Falcon+LLM%3A+Outperforming+Curated+Corpora+with+Web+Data%2C+and+Web+Data+Only 14. Quantifying Memorization Across Neural Language Models / broader Common Crawl representativeness and bias studies (e.g., work on Common Crawl's domain and language skew) — Various (Common Crawl bias/representativeness literature, e.g. work by Luccioni & Viviano on Common Crawl content quality, and follow-on studies), 2021-2023 https://scholar.google.com/scholar?q=Quantifying+Memorization+Across+Neural+Language+Models+%2F+broader+Common+Crawl+representativeness+and+bias+studies+%28e.g.%2C+work+on+Common+Crawl%27s+domain+and+language+skew%29 15. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research — Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, et al. (Allen Institute for AI), 2024 https://scholar.google.com/scholar?q=Dolma%3A+an+Open+Corpus+of+Three+Trillion+Tokens+for+Language+Model+Pretraining+Research 16. AI models collapse when trained on recursively generated data — I. Shumailov, Z. Shumaylov, Y. Zhao, N. Papernot, R. Anderson, Y. Gal, 2024 https://scholar.google.com/scholar?q=AI+models+collapse+when+trained+on+recursively+generated+data 17. Can AI-generated text be reliably detected? Stress testing AI text detectors under various attacks — V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, S. Feizi, 2025 https://scholar.google.com/scholar?q=Can+AI-generated+text+be+reliably+detected%3F+Stress+testing+AI+text+detectors+under+various+attacks 18. RAID: A shared benchmark for robust evaluation of machine-generated text detectors — L. Dugan, A. Hwang, F. Trhlik, A. Zhu, J. M. Ludan, H. Xu, D. Ippolito, C. Callison-Burch, 2024 https://scholar.google.com/scholar?q=RAID%3A+A+shared+benchmark+for+robust+evaluation+of+machine-generated+text+detectors 19. When incentives backfire, data stops being human — S. Santy, P. Bhattacharya, M. H. Ribeiro, K. Allen, S. Oh, 2025 https://scholar.google.com/scholar?q=When+incentives+backfire%2C+data+stops+being+human 20. Longitudinal sampling of URLs from the Wayback Machine — K. Garg, S. Alam, D. Ayala, M. Graham, M. C. Weigle, M. L. Nelson, 2025 https://scholar.google.com/scholar?q=Longitudinal+sampling+of+URLs+from+the+Wayback+Machine 21. Verbalized sampling: How to mitigate mode collapse and unlock LLM diversity — J. Zhang, S. Yu, D. Chong, A. Sicilia, M. R. Tomz, C. D. Manning, W. Shi, 2025 https://scholar.google.com/scholar?q=Verbalized+sampling%3A+How+to+mitigate+mode+collapse+and+unlock+LLM+diversity Interactive Visualization: AI-Generated Text and the Death of the Open Web
-
692
Approaching Shannon Bound: Lossless LLM Weight Compression
This episode explores "Approaching Shannon Bound with Lossless LLM Weight Compression," which argues that model weights stored in formats like bf16 carry far less real information than their bit-width implies—entropy measurements across six models and seven numeric formats show gaps of several bits per weight that can be recovered without any change to the underlying values. The discussion covers why memory capacity and bandwidth, not raw compute, are the real bottleneck in GPU inference, and why generic compressors like gzip fail on IEEE-754 floating point data. The hosts dig into Asymmetric Numeral Systems (ANS) as the key engineering breakthrough, since it decodes fast enough and in a tile-parallel enough fashion to run inside a live GPU kernel without becoming a new bottleneck itself. Listeners interested in the intersection of information theory and practical LLM serving will find the walkthrough of how lossless compression differs fundamentally from quantization methods like int4 or AWQ particularly compelling. Sources: 1. Approaching Shannon Bound with Lossless LLM Weight Compression — Hongshi Tan, Yao Chen, Gustavo Alonso, Weng-Fai Wong, Bingsheng He, 2026 http://arxiv.org/abs/2606.15789 2. 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float — Tianyi Zhang, Yang Sui, Shaochen (Henry) Zhong, et al., 2025 https://scholar.google.com/scholar?q=70%25+Size%2C+100%25+Accuracy%3A+Lossless+LLM+Compression+for+Efficient+GPU+Inference+via+Dynamic-Length+Float 3. NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks — Yongchang Hao, Yanshuai Cao, Lili Mou, 2024 https://scholar.google.com/scholar?q=NeuZip%3A+Memory-Efficient+Training+and+Inference+with+Dynamic+Compression+of+Neural+Networks 4. ZipNN: Lossless Compression for AI Models — Moshik Hershcovitch, Andrew Wood, Leshem Choshen, et al. (IBM Research), 2024 https://scholar.google.com/scholar?q=ZipNN%3A+Lossless+Compression+for+AI+Models 5. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding — Song Han, Huizi Mao, William J. Dally, 2016 https://scholar.google.com/scholar?q=Deep+Compression%3A+Compressing+Deep+Neural+Networks+with+Pruning%2C+Trained+Quantization+and+Huffman+Coding 6. Asymmetric Numeral Systems: Entropy Coding Combining Speed of Huffman Coding with Compression Rate of Arithmetic Coding — Jarek Duda, 2013 https://scholar.google.com/scholar?q=Asymmetric+Numeral+Systems%3A+Entropy+Coding+Combining+Speed+of+Huffman+Coding+with+Compression+Rate+of+Arithmetic+Coding 7. The Use of Asymmetric Numeral Systems as an Accurate Replacement for Huffman Coding — Jarek Duda, Khalid Tahboub, Neeraj J. Gadgil, Edward J. Delp, 2015 https://scholar.google.com/scholar?q=The+Use+of+Asymmetric+Numeral+Systems+as+an+Accurate+Replacement+for+Huffman+Coding 8. Variational Image Compression with a Scale Hyperprior — Johannes Ballé, David Minnen, Saurabh Singh, Sung Jin Hwang, Nick Johnston, 2018 https://scholar.google.com/scholar?q=Variational+Image+Compression+with+a+Scale+Hyperprior 9. Zstandard Compression and the application/zstd Media Type (RFC 8878) — Yann Collet, Murray Kucherawy (eds.), 2020 https://scholar.google.com/scholar?q=Zstandard+Compression+and+the+application%2Fzstd+Media+Type+%28RFC+8878%29 10. GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLMs — Hao Kang, Qingru Zhang, Souvik Kundu, Geonhwa Jeong, Zaoxing Liu, Tushar Krishna, Tuo Zhao, 2024 https://scholar.google.com/scholar?q=GEAR%3A+An+Efficient+KV+Cache+Compression+Recipe+for+Near-Lossless+Generative+Inference+of+LLMs 11. S-LoRA: Serving Thousands of Concurrent LoRA Adapters — Ying Sheng, Shiyi Cao, Dacheng Li, Coleman Hooper, Nicholas Lee, Shuo Yang, Christopher Chou, Banghua Zhu, Lianmin Zheng, Kurt Keutzer, Joseph E. Gonzalez, Ion Stoica, 2023 https://scholar.google.com/scholar?q=S-LoRA%3A+Serving+Thousands+of+Concurrent+LoRA+Adapters 12. QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving — Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, Song Han, 2024 https://scholar.google.com/scholar?q=QServe%3A+W4A8KV4+Quantization+and+System+Co-design+for+Efficient+LLM+Serving 13. Efficient LLM Inference: Bandwidth, Compute, Synchronization, and Capacity are All You Need — M. Davies, N. Crago, K. Sankaralingam, C. Kozyrakis, 2025 https://scholar.google.com/scholar?q=Efficient+LLM+Inference%3A+Bandwidth%2C+Compute%2C+Synchronization%2C+and+Capacity+are+All+You+Need Interactive Visualization: Approaching Shannon Bound: Lossless LLM Weight Compression
-
691
Test-Time Adaptation Through Entropy Minimization
This episode explores Tent, a 2021 method for adapting a frozen classifier to shifted test data using only entropy minimization on unlabeled target inputs — no retraining, no labels, and no access to the original source dataset. It contrasts this "fully test-time adaptation" setting against classical domain adaptation, which still requires the source data on hand during adjustment, a constraint that's often impractical for vendors shipping models under privacy or bandwidth limits. The discussion digs into the mechanism: Tent re-estimates BatchNorm statistics on incoming test batches and tunes only the tiny per-channel scale-and-shift parameters (under 1% of the network), repurposing existing training infrastructure for adaptation. The hosts also interrogate the core intuition — that confident predictions tend to be correct — pressing on whether that assumption holds when a model's decision boundaries are already unreliable, without fully resolving the tension before turning to results. Listeners interested in low-cost deployment fixes for distribution shift, or skeptical of self-referential confidence-based methods, will find the back-and-forth pushback especially engaging. Sources: 1. Tent: Fully Test-time Adaptation by Entropy Minimization — Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, Trevor Darrell, 2020 http://arxiv.org/abs/2006.10726 2. Semi-Supervised Learning by Entropy Minimization — Yves Grandvalet, Yoshua Bengio, 2004 https://scholar.google.com/scholar?q=Semi-Supervised+Learning+by+Entropy+Minimization 3. A DIRT-T Approach to Unsupervised Domain Adaptation — Rui Shu, Hung Bui, Hirokazu Narui, Stefano Ermon, 2018 https://scholar.google.com/scholar?q=A+DIRT-T+Approach+to+Unsupervised+Domain+Adaptation 4. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts — Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, Moritz Hardt, 2020 https://scholar.google.com/scholar?q=Test-Time+Training+with+Self-Supervision+for+Generalization+under+Distribution+Shifts 5. Unsupervised Domain Adaptation by Backpropagation — Yaroslav Ganin, Victor Lempitsky, 2015 https://scholar.google.com/scholar?q=Unsupervised+Domain+Adaptation+by+Backpropagation 6. Learning from Synthetic Data: Addressing Domain Shift for Semantic Segmentation — Swami Sankaranarayanan, Yogesh Balaji, Arpit Jain, Ser Nam Lim, Rama Chellappa, 2018 https://scholar.google.com/scholar?q=Learning+from+Synthetic+Data%3A+Addressing+Domain+Shift+for+Semantic+Segmentation 7. Revisiting Batch Normalization For Practical Domain Adaptation — Yanghao Li, Naiyan Wang, Jianping Shi, Jiaying Liu, Xiaodi Hou, 2016 https://scholar.google.com/scholar?q=Revisiting+Batch+Normalization+For+Practical+Domain+Adaptation 8. CyCADA: Cycle-Consistent Adversarial Domain Adaptation — Judy Hoffman, Eric Tzeng, Taesung Park, Jun-Yan Zhu, Phillip Isola, Kate Saenko, Alexei A. Efros, Trevor Darrell, 2018 https://scholar.google.com/scholar?q=CyCADA%3A+Cycle-Consistent+Adversarial+Domain+Adaptation 9. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift — Sergey Ioffe, Christian Szegedy, 2015 https://scholar.google.com/scholar?q=Batch+Normalization%3A+Accelerating+Deep+Network+Training+by+Reducing+Internal+Covariate+Shift 10. Evaluating Prediction-Time Batch Normalization for Robustness to Covariate Shift — Zachary Nado, Shreyas Padhy, D. Sculley, Alexander D'Amour, Balaji Lakshminarayanan, Jasper Snoek, 2020 https://scholar.google.com/scholar?q=Evaluating+Prediction-Time+Batch+Normalization+for+Robustness+to+Covariate+Shift 11. Improving Robustness Against Common Corruptions by Covariate Shift Adaptation — Steffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann, Wieland Brendel, Matthias Bethge, 2020 https://scholar.google.com/scholar?q=Improving+Robustness+Against+Common+Corruptions+by+Covariate+Shift+Adaptation 12. Test-Time Training for Out-of-Distribution Generalization — Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, Moritz Hardt, 2019 https://scholar.google.com/scholar?q=Test-Time+Training+for+Out-of-Distribution+Generalization 13. Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain Adaptation (SHOT) — Jian Liang, Dapeng Hu, Jiashi Feng, 2020 https://scholar.google.com/scholar?q=Do+We+Really+Need+to+Access+the+Source+Data%3F+Source+Hypothesis+Transfer+for+Unsupervised+Domain+Adaptation+%28SHOT%29 14. Do CIFAR-10 Classifiers Generalize to CIFAR-10? / Do ImageNet Classifiers Generalize to ImageNet? — Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, Vaishaal Shankar, 2018 / 2019 https://scholar.google.com/scholar?q=Do+CIFAR-10+Classifiers+Generalize+to+CIFAR-10%3F+%2F+Do+ImageNet+Classifiers+Generalize+to+ImageNet%3F 15. A Simple Way to Make Neural Networks Robust Against Diverse Image Corruptions (ANT) — Evgenia Rusak, Lukas Schott, Roland S. Zimmermann, Julian Bitterwolf, Oliver Bringmann, Matthias Bethge, Wieland Brendel, 2020 https://scholar.google.com/scholar?q=A+Simple+Way+to+Make+Neural+Networks+Robust+Against+Diverse+Image+Corruptions+%28ANT%29 Interactive Visualization: Test-Time Adaptation Through Entropy Minimization
-
690
SOAP: Stabilizing Shampoo's Second-Order Optimizer with Adam
This episode explores SOAP, a new optimizer from a Harvard/Kempner Institute team that fuses Shampoo's second-order preconditioning with Adam's update mechanics. The hosts trace the lineage from Adagrad's mathematically ideal but computationally infeasible full preconditioner matrix, through Adam's cheap diagonal approximation, to Shampoo's middle-ground Kronecker-product approach using two smaller per-dimension preconditioners. The core theoretical result discussed is a proof that Shampoo run with the one-half power is mathematically equivalent to running Adafactor inside the eigenbasis Shampoo's own preconditioner defines — which motivates simply swapping in full Adam within that same rotated basis, adding just one new hyperparameter (preconditioning frequency) over standard AdamW. The discussion highlights the paper's striking efficiency claims — over 40% fewer training iterations and 35% less wall-clock time versus AdamW, and roughly 20% better than Shampoo itself — while noting these numbers deserve scrutiny given real-world context like Shampoo's AlgoPerf benchmark win and its use in training Gemini 1.5 Flash. Listeners interested in the mechanics behind large-scale training efficiency will get a clear breakdown of why optimizer choice translates directly into cluster-scale compute costs and calendar time. Sources: 1. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, Mujin Kwun, Itai Shapira, David Brandfonbrener, Lucas Janson, Sham Kakade, 2024 http://arxiv.org/abs/2409.11321 2. Muon: An optimizer for hidden layers in neural networks — Keller Jordan et al., 2024 https://scholar.google.com/scholar?q=Muon%3A+An+optimizer+for+hidden+layers+in+neural+networks 3. 4-bit Shampoo for Memory-Efficient Network Training — Sike Wang, Jia Li, Pan Zhou, Hua Huang, 2024 https://scholar.google.com/scholar?q=4-bit+Shampoo+for+Memory-Efficient+Network+Training 4. Combining axes preconditioners through Kronecker approximation for deep learning — Sai Surya Duvvuri, Fnu Devvrit, Rohan Anil, Cho-Jui Hsieh, Inderjit S. Dhillon, 2024 https://scholar.google.com/scholar?q=Combining+axes+preconditioners+through+Kronecker+approximation+for+deep+learning 5. No train no gain: Revisiting efficient training algorithms for transformer-based language models — Jean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini, Matt J. Kusner, 2023 https://scholar.google.com/scholar?q=No+train+no+gain%3A+Revisiting+efficient+training+algorithms+for+transformer-based+language+models Interactive Visualization: SOAP: Stabilizing Shampoo's Second-Order Optimizer with Adam
-
689
Distributed Shampoo: Making Second-Order Optimization Practical at Scale
This episode explores a distributed-systems paper from Meta, Mila/University of Montreal, and NVIDIA researchers that tackles a long-standing tradeoff in neural network training: second-order-style optimizers like Shampoo converge better than the ubiquitous diagonal methods (AdaGrad, RMSProp, Adam), but were assumed too computationally expensive to use at scale. The discussion traces the lineage from diagonal AdaGrad through the impractical "full-matrix" AdaGrad to Shampoo's key innovation — approximating each layer's preconditioner as a Kronecker product of two much smaller matrices, drawing a parallel to Martens and Grosse's independently-derived KFAC method. The hosts debate whether Shampoo counts as a true second-order method or something distinct rooted in online convex optimization theory, and highlight the paper's headline systems result: a distributed PyTorch implementation that keeps the wall-clock overhead of this matrix-based optimizer to roughly 10% per step. Listeners interested in the practical engineering behind making theoretically superior optimizers actually usable at billion-parameter scale will find the breakdown of the memory and compute tradeoffs especially compelling. Sources: 1. A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale — Hao-Jun Michael Shi, Tsung-Hsien Lee, Shintaro Iwasaki, Jose Gallego-Posada, Zhijing Li, Kaushik Rangadurai, Dheevatsa Mudigere, Michael Rabbat, 2023 http://arxiv.org/abs/2309.06497 2. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization — John Duchi, Elad Hazan, Yoram Singer, 2011 https://scholar.google.com/scholar?q=Adaptive+Subgradient+Methods+for+Online+Learning+and+Stochastic+Optimization 3. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization 4. Scalable Second Order Optimization for Deep Learning — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer (Google), 2020 https://scholar.google.com/scholar?q=Scalable+Second+Order+Optimization+for+Deep+Learning 5. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature 6. Optimizing Neural Networks with Kronecker-factored Approximate Curvature (K-FAC) — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature+%28K-FAC%29 7. On the Factory Floor: ML Engineering for Industrial-Scale Ads Recommendation Models — Rohan Anil, Sandra Gadanho, Da Huang, et al., 2022 https://scholar.google.com/scholar?q=On+the+Factory+Floor%3A+ML+Engineering+for+Industrial-Scale+Ads+Recommendation+Models 8. Adafactor: Adaptive Learning Rates with Sublinear Memory Cost — Noam Shazeer, Mitchell Stern, 2018 https://scholar.google.com/scholar?q=Adafactor%3A+Adaptive+Learning+Rates+with+Sublinear+Memory+Cost 9. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, et al., 2024 https://scholar.google.com/scholar?q=SOAP%3A+Improving+and+Stabilizing+Shampoo+using+Adam Interactive Visualization: Distributed Shampoo: Making Second-Order Optimization Practical at Scale
-
688
Scalable Second-Order Optimization: Shampoo at Scale
This episode explores the paper "Scalable Second Order Optimization for Deep Learning" and its introduction of Distributed Shampoo, a Kronecker-factored second-order optimizer that cuts training steps in half compared to a well-tuned Adam baseline on WMT'14 English-to-French translation, with further gains shown on BERT, Criteo click-through-rate modeling, and ResNet-50. The discussion traces the lineage from Newton's method and full-matrix AdaGrad's prohibitive O(N²)/O(N³) costs through Shampoo's Kronecker-product approximation, which replaces one massive preconditioner with smaller per-dimension matrices to make second-order optimization tractable at scale. Rival factored approaches, K-FAC and K-BFGS, are positioned as points of comparison throughout the paper. The conversation is aimed at listeners who use Adam daily but have never unpacked the mechanical distinction between first-order and second-order optimization, or the meaning of "preconditioning" itself. It's a compelling listen because second-order methods have long been dismissed as theoretically superior but practically unscalable, and this paper demonstrates real wall-clock wins across four distinct production-scale workloads rather than a single cherry-picked benchmark. Sources: 1. Scalable Second Order Optimization for Deep Learning — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer, 2020 http://arxiv.org/abs/2002.09018 2. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization 3. Optimizing Neural Networks with Kronecker-factored Approximate Curvature (K-FAC) — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature+%28K-FAC%29 4. Adam: A Method for Stochastic Optimization — Diederik P. Kingma, Jimmy Ba, 2014 https://scholar.google.com/scholar?q=Adam%3A+A+Method+for+Stochastic+Optimization 5. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization (AdaGrad) — John Duchi, Elad Hazan, Yoram Singer, 2011 https://scholar.google.com/scholar?q=Adaptive+Subgradient+Methods+for+Online+Learning+and+Stochastic+Optimization+%28AdaGrad%29 6. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — J. Martens, R. Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature 7. Distributed Second-Order Optimization using Kronecker-Factored Approximations — J. Ba, J. Martens, R. Grosse, 2017 https://scholar.google.com/scholar?q=Distributed+Second-Order+Optimization+using+Kronecker-Factored+Approximations 8. Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks — K. Osawa, Y. Tsuji, Y. Ueno, A. Naruse, R. Yokota, S. Matsuoka, 2019 https://scholar.google.com/scholar?q=Large-Scale+Distributed+Second-Order+Optimization+Using+Kronecker-Factored+Approximate+Curvature+for+Deep+Convolutional+Neural+Networks 9. Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes — Y. You, J. Li, S. Reddi, et al. (LAMB), 2019 https://scholar.google.com/scholar?q=Large+Batch+Optimization+for+Deep+Learning%3A+Training+BERT+in+76+Minutes 10. A Large Batch Optimizer Reality Check: Traditional, Generic Optimizers Suffice Across Batch Sizes — Z. Nado, J. Gilmer, C. Shallue, R. Anil, G. Dahl, 2021 https://scholar.google.com/scholar?q=A+Large+Batch+Optimizer+Reality+Check%3A+Traditional%2C+Generic+Optimizers+Suffice+Across+Batch+Sizes 11. Limitations of the Empirical Fisher Approximation for Natural Gradient Descent — F. Kunstner, P. Hennig, L. Balles, 2019 https://scholar.google.com/scholar?q=Limitations+of+the+Empirical+Fisher+Approximation+for+Natural+Gradient+Descent Interactive Visualization: Scalable Second-Order Optimization: Shampoo at Scale
-
687
Preconditioned Optimization Without the Full Matrix Cost
This episode explores Shampoo, a 2018 optimization algorithm from Google Brain and Princeton that brings second-order curvature information to neural network training without the prohibitive cost of a full preconditioner. The discussion traces the lineage from Newton's method through AdaGrad's diagonal and full-matrix variants, explaining how Shampoo exploits the natural tensor structure of neural network weights — keeping separate small preconditioners per dimension and combining them via Kronecker products rather than materializing an impossibly large matrix. The hosts debate how much weight to put on the algorithm's convergence proofs, which rely on online convex optimization theory even though real network training is highly non-convex, concluding that the theory offers a sanity check rather than a guarantee and that empirical performance does the real work of justification. Along the way, they place Shampoo alongside K-FAC as one of the major structure-aware curvature approximations in the optimization literature. Listeners interested in the tradeoffs between cheap first-order methods and expensive second-order ones will find a clear walkthrough of how Shampoo threads that needle. Sources: 1. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 http://arxiv.org/abs/1802.09568 2. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization — John Duchi, Elad Hazan, Yoram Singer, 2011 https://scholar.google.com/scholar?q=Adaptive+Subgradient+Methods+for+Online+Learning+and+Stochastic+Optimization 3. Scalable Second Order Optimization for Deep Learning (Distributed Shampoo) — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer, 2021 https://scholar.google.com/scholar?q=Scalable+Second+Order+Optimization+for+Deep+Learning+%28Distributed+Shampoo%29 4. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira, David Brandfonbrener, Sham Kakade, 2024 https://scholar.google.com/scholar?q=SOAP%3A+Improving+and+Stabilizing+Shampoo+using+Adam 5. Muon: An optimizer for hidden layers in neural networks — Keller Jordan et al., 2024 https://scholar.google.com/scholar?q=Muon%3A+An+optimizer+for+hidden+layers+in+neural+networks 6. Adam: A Method for Stochastic Optimization — Diederik Kingma, Jimmy Ba, 2015 https://scholar.google.com/scholar?q=Adam%3A+A+Method+for+Stochastic+Optimization 7. Decoupled Weight Decay Regularization — Ilya Loshchilov, Frank Hutter, 2017 https://scholar.google.com/scholar?q=Decoupled+Weight+Decay+Regularization 8. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature 9. A Stochastic Quasi-Newton Method for Large-Scale Optimization — Richard Byrd, Samantha Hansen, Jorge Nocedal, Yoran Singer, 2016 https://scholar.google.com/scholar?q=A+Stochastic+Quasi-Newton+Method+for+Large-Scale+Optimization 10. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization 11. Online Convex Programming and Generalized Infinitesimal Gradient Ascent — Martin Zinkevich, 2003 https://scholar.google.com/scholar?q=Online+Convex+Programming+and+Generalized+Infinitesimal+Gradient+Ascent 12. Online Learning and Online Convex Optimization — Shai Shalev-Shwartz, 2012 https://scholar.google.com/scholar?q=Online+Learning+and+Online+Convex+Optimization 13. Attention Is All You Need — A. Vaswani, N. Shazeer, N. Parmar, et al., 2017 https://scholar.google.com/scholar?q=Attention+Is+All+You+Need 14. A Unified Approach to Adaptive Regularization in Online and Stochastic Optimization — V. Gupta, T. Koren, Y. Singer, 2017 https://scholar.google.com/scholar?q=A+Unified+Approach+to+Adaptive+Regularization+in+Online+and+Stochastic+Optimization 15. Understanding Deep Learning Requires Rethinking Generalization — C. Zhang, S. Bengio, M. Hardt, B. Recht, O. Vinyals, 2017 https://scholar.google.com/scholar?q=Understanding+Deep+Learning+Requires+Rethinking+Generalization Interactive Visualization: Preconditioned Optimization Without the Full Matrix Cost
-
686
Naive Test-Time Adaptation Destabilizes LLM Predictions
This episode explores a paper proposing SCALENET, a hypernetwork-based fix for unsupervised test-time adaptation (TTA) in large language models. It examines why naive per-prompt gradient updates are unstable — a 70-billion-parameter Llama model's negative log-likelihood balloons from 2.21 to 11.49 after just five adaptation steps — and traces the problem to high-variance single-sample gradients that can't average out the way batch training does. The discussion covers the constrained "adapt-and-reset" setup used in real deployment, where models take a few unsupervised gradient steps on LoRA attention matrices per prompt before discarding the update, and explains why a single global learning rate can't work when small rates do nothing and large ones destroy the model. Listeners interested in the mechanics of on-the-fly model adaptation, LoRA-based efficient tuning, and the control-theory-like challenge of stabilizing per-layer, per-step learning rates will find the breakdown of the failure modes and the proposed hypernetwork solution especially compelling. Sources: 1. Unsupervised Layer-Wise Dynamic Test Time Adaptation for LLMs — Longhuan Xu, Cunjian Chen, Feng Yin, 2026 http://arxiv.org/abs/2602.09719 2. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts — Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei Efros, Moritz Hardt, 2020 https://scholar.google.com/scholar?q=Test-Time+Training+with+Self-Supervision+for+Generalization+under+Distribution+Shifts 3. Tent: Fully Test-Time Adaptation by Entropy Minimization — Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, Trevor Darrell, 2021 https://scholar.google.com/scholar?q=Tent%3A+Fully+Test-Time+Adaptation+by+Entropy+Minimization 4. Test-Time Training on Nearest Neighbors for Large Language Models — Moritz Hardt, Yu Sun, 2024 https://scholar.google.com/scholar?q=Test-Time+Training+on+Nearest+Neighbors+for+Large+Language+Models 5. The Surprising Effectiveness of Test-Time Training for Abstract Reasoning — Ekin Akyürek, Mehul Damani, Linlu Qiu, Han Guo, Yoon Kim, Jacob Andreas, 2024 https://scholar.google.com/scholar?q=The+Surprising+Effectiveness+of+Test-Time+Training+for+Abstract+Reasoning 6. Test-time Learning for Large Language Models — Hu, J., Zhang, Z., Chen, G., Wen, X., Shuai, C., Luo, W., Xiao, B., Li, Y., Tan, M., 2025 https://scholar.google.com/scholar?q=Test-time+Learning+for+Large+Language+Models 7. SLOT: Sample-specific Language Model Optimization at Test-time — Hu, Y., Zhang, X., Fang, X., Chen, Z., Wang, X., Zhang, H., Qi, G., 2025 https://scholar.google.com/scholar?q=SLOT%3A+Sample-specific+Language+Model+Optimization+at+Test-time 8. COME: Test-time Adaption by Conservatively Minimizing Entropy — Zhang, Q., Bian, Y., Kong, X., Zhao, P., Zhang, C., 2024 https://scholar.google.com/scholar?q=COME%3A+Test-time+Adaption+by+Conservatively+Minimizing+Entropy 9. Revisiting Dynamic Evaluation: Online Adaptation for Large Language Models — Rannen-Triki, A., Bornschein, J., Pascanu, R., Hutter, M., et al., 2024 https://scholar.google.com/scholar?q=Revisiting+Dynamic+Evaluation%3A+Online+Adaptation+for+Large+Language+Models 10. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks (MAML) — Finn, C., Abbeel, P., Levine, S., 2017 https://scholar.google.com/scholar?q=Model-Agnostic+Meta-Learning+for+Fast+Adaptation+of+Deep+Networks+%28MAML%29 Interactive Visualization: Naive Test-Time Adaptation Destabilizes LLM Predictions
-
685
Kohonen's 1972 Correlation Matrix Memory, Decades Before Attention
This episode revisits Teuvo Kohonen's 1972 paper "Correlation Matrix Memories," which reframes associative memory as a hardware fault-tolerance problem rather than a representation-learning one. Kohonen builds a memory from outer-product sums of key and data vectors, then shows mathematically how much recall quality degrades when connections are randomly dropped (an "incomplete" correlation matrix memory) rather than fully wired. The discussion traces the paper's lineage against optical holography models and Steinbuch's Lernmatrix, and unpacks concepts like crosstalk and graceful degradation as information gets smeared additively across the matrix instead of stored in one fragile spot. A tangent draws — and partly disputes — a comparison between Kohonen's outer-product accumulation and the mechanics underlying modern attention, debating whether the resemblance is structural or purely coincidental given the total absence of learning or gradients in the original scheme. Listeners interested in the deep history of neural memory models and how old hardware constraints shaped ideas that echo in today's architectures will find plenty to chew on. Sources: 1. Kohonen's 1972 Correlation Matrix Memory, Decades Before Attention https://lucidar.me/fr/neural-networks/files/1972-correlation-matrix-memories.pdf 2. Neural networks and physical systems with emergent collective computational abilities — John J. Hopfield, 1982 https://scholar.google.com/scholar?q=Neural+networks+and+physical+systems+with+emergent+collective+computational+abilities 3. Non-Holographic Associative Memory — David Willshaw, O. P. Buneman, H. Christopher Longuet-Higgins, 1969 https://scholar.google.com/scholar?q=Non-Holographic+Associative+Memory 4. Hopfield Networks is All You Need — Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, et al., 2020 https://scholar.google.com/scholar?q=Hopfield+Networks+is+All+You+Need 5. Linear Transformers Are Secretly Fast Weight Programmers — Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber, 2021 https://scholar.google.com/scholar?q=Linear+Transformers+Are+Secretly+Fast+Weight+Programmers 6. A Simple Neural Network Generating an Interactive Memory — James A. Anderson, 1972 https://scholar.google.com/scholar?q=A+Simple+Neural+Network+Generating+an+Interactive+Memory 7. Representation of Associated Data by Matrix Operators — Teuvo Kohonen, Matti Ruohonen, 1973 https://scholar.google.com/scholar?q=Representation+of+Associated+Data+by+Matrix+Operators 8. Sparse Distributed Memory — Pentti Kanerva, 1988 https://scholar.google.com/scholar?q=Sparse+Distributed+Memory 9. Die Lernmatrix — K. Steinbuch, 1961 https://scholar.google.com/scholar?q=Die+Lernmatrix 10. Associative holographic memories — D. Gabor, 1969 https://scholar.google.com/scholar?q=Associative+holographic+memories 11. A class of randomly organized associative memories — T. Kohonen, 1971 https://scholar.google.com/scholar?q=A+class+of+randomly+organized+associative+memories Interactive Visualization: Kohonen's 1972 Correlation Matrix Memory, Decades Before Attention
-
684
Model-Agnostic Meta-Learning for Fast Task Adaptation
This episode explores Model-Agnostic Meta-Learning (MAML), the 2017 approach from Chelsea Finn, Pieter Abbeel, and Sergey Levine that trains a single, architecture-agnostic initialization capable of fast adaptation across image classification, regression, and reinforcement learning. Rather than learning a task-specific update rule like earlier recurrent meta-learners, MAML optimizes the starting weights themselves so that a few steps of ordinary gradient descent adapt them well to a brand-new task from minimal data, tested through Omniglot and MiniImagenet few-shot classification, sinusoid regression, and MuJoCo/2D navigation RL. The discussion breaks down the inner-loop/outer-loop structure, the second-order gradient-through-gradient math (Hessian-vector products) needed to backpropagate through the adaptation step, and how finite-difference approximations sidestep the third-derivative problem when TRPO is used as the RL meta-optimizer. Listeners get a clear walkthrough of N-way K-shot learning and why one image per class is such an extreme test of generalization, plus a grounded comparison to the more familiar pretrain-then-fine-tune workflow. It's a good listen for anyone curious how a deceptively simple idea — learn to be easy to fine-tune — unified meta-learning across problem types that previously required separate specialized systems. Sources: 1. Model-Agnostic Meta-Learning for Fast Task Adaptation https://proceedings.mlr.press/v70/finn17a/finn17a.pdf 2. Optimization as a Model for Few-Shot Learning — Sachin Ravi, Hugo Larochelle, 2017 https://scholar.google.com/scholar?q=Optimization+as+a+Model+for+Few-Shot+Learning 3. Meta-Learning with Memory-Augmented Neural Networks — Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, Timothy Lillicrap, 2016 https://scholar.google.com/scholar?q=Meta-Learning+with+Memory-Augmented+Neural+Networks 4. Matching Networks for One Shot Learning — Oriol Vinyals, Charles Blundell, Tim Lillicrap, Daan Wierstra, 2016 https://scholar.google.com/scholar?q=Matching+Networks+for+One+Shot+Learning 5. Learning to Learn by Gradient Descent by Gradient Descent — Marcin Andrychowicz et al., 2016 https://scholar.google.com/scholar?q=Learning+to+Learn+by+Gradient+Descent+by+Gradient+Descent 6. Trust Region Policy Optimization — John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, Philipp Moritz, 2015 https://scholar.google.com/scholar?q=Trust+Region+Policy+Optimization 7. RL2: Fast Reinforcement Learning via Slow Reinforcement Learning — Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, Pieter Abbeel, 2016 https://scholar.google.com/scholar?q=RL2%3A+Fast+Reinforcement+Learning+via+Slow+Reinforcement+Learning Interactive Visualization: Model-Agnostic Meta-Learning for Fast Task Adaptation
-
683
In-Place Test-Time Training Turns Fast Weights Into Online Memory
This episode explores a new test-time training method called In-Place TTT, which repurposes the down-projection matrix inside a model's existing gated MLP as adaptable "fast weights," letting a pretrained model keep learning during inference without any architectural changes. A key innovation is replacing the reconstruction-style training target used in prior TTT approaches with an LM-aligned target built from a causal convolution over token embeddings, which the authors prove (via an induction-head theorem) actually raises the probability of the correct next token. The discussion covers how a context-parallel scan preserves causality while enabling parallel computation of these updates, and walks through benchmark results showing the method trailing a baseline at short context but pulling substantially ahead as sequence length grows, tested across Qwen3-4B, LLaMA-3.1-8B, and Qwen3-14B. The hosts also dig into an ablation showing that mid-sized chunk sizes outperform larger ones — a counterintuitive result tied to how often the fast weights get to update rather than raw parallelism — plus efficiency data showing the approach barely affects throughput or memory. It's a concrete look at how far you can push adaptive inference-time learning while reusing a model's own existing structure. Sources: 1. In-Place Test-Time Training — Guhao Feng, Shengjie Luo, Kai Hua, Ge Zhang, Di He, Wenhao Huang, Tianle Cai, 2026 http://arxiv.org/abs/2604.06169 2. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun, Xinhao Li, Karan Dalal, et al., 2024 https://scholar.google.com/scholar?q=Learning+to+%28Learn+at+Test+Time%29%3A+RNNs+with+Expressive+Hidden+States 3. Test-Time Training Done Right (LaCT) — Tianyuan Zhang, Sai Bi, Yicong Hong, et al., 2025 https://scholar.google.com/scholar?q=Test-Time+Training+Done+Right+%28LaCT%29 4. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2024 https://scholar.google.com/scholar?q=Titans%3A+Learning+to+Memorize+at+Test+Time 5. Transformer Feed-Forward Layers Are Key-Value Memories — Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy, 2020 https://scholar.google.com/scholar?q=Transformer+Feed-Forward+Layers+Are+Key-Value+Memories 6. Locating and Editing Factual Associations in GPT (ROME) — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022 https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT+%28ROME%29 7. LoRA: Low-Rank Adaptation of Large Language Models — Edward Hu, Yelong Shen, Phillip Wallis, et al., 2022 https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models Interactive Visualization: In-Place Test-Time Training Turns Fast Weights Into Online Memory
-
682
Hal: Memory-Augmented LLM Agents Still Hit Continual Learning's Wall
This episode explores a paper examining what happens to continual learning problems when LLM agents shift from parametric updates to memory-augmented architectures. Rather than accepting the industry assumption that external memory sidesteps catastrophic forgetting entirely, the researchers run classic continual-learning protocols on memory-based agents and find the same core problem resurfaces in a new form — shifting from parameter capacity to context-window retrieval capacity. They identify three specific failure modes: retrieval pollution (irrelevant memories crowding the prompt), context competition (useful memories getting displaced by other retrieved items), and memory dilution (relevant material becoming harder to surface as the memory store grows). The discussion traces this argument against the history of catastrophic forgetting and prior mitigation techniques like Elastic Weight Consolidation and Gradient Episodic Memory, then explains how the paper reframes the stability-plasticity dilemma for retrieval-based systems. Listeners interested in agent design, RAG architectures, or the assumptions underlying memory-augmented LLMs will find the paper's reframing — that memory doesn't eliminate the bottleneck, it just relocates it — a useful corrective to a widely repeated industry pitch. Sources: 1. Hal: Memory-Augmented LLM Agents Still Hit Continual Learning's Wall https://arxiv.org/pdf/2604.27003 2. A-Mem: Agentic Memory for LLM Agents — Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, Yongfeng Zhang, 2025 https://scholar.google.com/scholar?q=A-Mem%3A+Agentic+Memory+for+LLM+Agents 3. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Vivian Fang, Shishir G. Patil, Kevin Lin, Sarah Wooders, Joseph E. Gonzalez, 2023 https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems 4. How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior — Zidi Xiong, Yuping Lin, Wenya Xie, Pengfei He, Zirui Liu, Jiliang Tang, Himabindu Lakkaraju, Zhen Xiang, 2025 https://scholar.google.com/scholar?q=How+Memory+Management+Impacts+LLM+Agents%3A+An+Empirical+Study+of+Experience-Following+Behavior 5. From RAG to Memory: Non-Parametric Continual Learning for Large Language Models — Bernal Jiménez Gutiérrez, Yiheng Shu, Weijian Qi, Sizhe Zhou, Yu Su, 2025 https://scholar.google.com/scholar?q=From+RAG+to+Memory%3A+Non-Parametric+Continual+Learning+for+Large+Language+Models 6. The Probabilistic Relevance Framework: BM25 and Beyond — Stephen Robertson, Hugo Zaragoza, 2009 https://scholar.google.com/scholar?q=The+Probabilistic+Relevance+Framework%3A+BM25+and+Beyond Interactive Visualization: Hal: Memory-Augmented LLM Agents Still Hit Continual Learning's Wall
-
681
Continual Learning in LLMs: Beyond Catastrophic Forgetting
This episode explores a survey on continual learning in large language models, examining how models can be updated after pretraining without the prohibitive cost of full retraining or the risk of catastrophic forgetting — the phenomenon where new training quietly degrades performance on tasks a model previously handled well. The discussion breaks down the problem across three distinct LLM training stages (pretraining, fine-tuning, and alignment) and maps them onto three classical mitigation strategies: rehearsal-based methods that replay old data, regularization-based methods that penalize changes to critical parameters, and architecture-based methods that add task-specific capacity like adapters or LoRA modules while freezing the rest. The hosts debate the survey's core organizational claim — that structuring the literature by mechanism rather than by application domain (medical, legal, financial) offers a more useful lens for practitioners trying to borrow a specific forgetting-mitigation technique. Listeners interested in the practical tradeoffs of keeping frontier models current — especially around data that can never legally enter a pretraining corpus, like medical or financial records — will find this a grounded framing of a problem every deployed LLM eventually faces. Sources: 1. Continual Learning in Large Language Models: Methods, Challenges, and Opportunities — Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin, 2026 http://arxiv.org/abs/2603.12658 2. Don't Stop Pretraining: Adapt Language Models to Domains and Tasks — Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, Noah A. Smith, 2020 https://scholar.google.com/scholar?q=Don%27t+Stop+Pretraining%3A+Adapt+Language+Models+to+Domains+and+Tasks 3. Simple and Scalable Strategies to Continually Pre-train Large Language Models — Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L. Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, Irina Rish, 2024 https://scholar.google.com/scholar?q=Simple+and+Scalable+Strategies+to+Continually+Pre-train+Large+Language+Models 4. LLaMA Pro: Progressive LLaMA with Block Expansion — Chengyue Wu, Yukang Gan, Yixiao Ge, Zeyu Lu, Jiahao Wang, Ye Feng, Ying Shan, Ping Luo, 2024 https://scholar.google.com/scholar?q=LLaMA+Pro%3A+Progressive+LLaMA+with+Block+Expansion 5. Code Llama: Open Foundation Models for Code — Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, and the Code Llama team at Meta AI, 2023 https://scholar.google.com/scholar?q=Code+Llama%3A+Open+Foundation+Models+for+Code 6. Editing Models with Task Arithmetic — Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, Ali Farhadi, 2023 https://scholar.google.com/scholar?q=Editing+Models+with+Task+Arithmetic 7. TIES-Merging: Resolving Interference When Merging Models — Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, Mohit Bansal, 2023 https://scholar.google.com/scholar?q=TIES-Merging%3A+Resolving+Interference+When+Merging+Models 8. Overcoming Catastrophic Forgetting in Neural Networks (EWC) — James Kirkpatrick et al., 2017 https://scholar.google.com/scholar?q=Overcoming+Catastrophic+Forgetting+in+Neural+Networks+%28EWC%29 Interactive Visualization: Continual Learning in LLMs: Beyond Catastrophic Forgetting
-
680
Learning, Fast and Slow: LLMs That Adapt Without Forgetting
This episode explores catastrophic forgetting and plasticity loss in RL-trained language models, and introduces "Fast-Slow Training," a method combining slow weight updates (RLVR) with fast in-context learning to address both. The hosts unpack the distinction between RLVR's automatic, verifiable rewards and traditional RLHF, then dig into two separate failure modes of pure RL post-training: models forgetting general competence while chasing a narrow reward signal, and a subtler loss of plasticity where updates leave models increasingly unable to absorb new tasks. Framing the two training channels as a System 1/System 2 split, the discussion centers on the paper's headline result — combining both channels reaches RL's peak accuracy with up to three times fewer samples, drifts up to seventy percent less from the base model, and preserves the capacity to learn subsequent tasks where pure RL stalls. Listeners interested in the mechanics and tradeoffs of continual learning in large language models will find a grounded walkthrough of why prompting alone hits a ceiling and why weight updates alone come with hidden costs. Sources: 1. Learning, Fast and Slow: Towards LLMs That Adapt Continually — Rishabh Tiwari, Kusha Sareen, Lakshya A Agrawal, Joseph E. Gonzalez, Matei Zaharia, Kurt Keutzer, Inderjit S Dhillon, Rishabh Agarwal, Devvrit Khatri, 2026 http://arxiv.org/abs/2605.12484 2. Loss of Plasticity in Deep Continual Learning — Shibhansh Dohare, J. Fernando Hernandez-Garcia, Qingfeng Lan, A. Rupam Mahmood, Richard S. Sutton, et al., 2024 (Nature; preprint circulated as 'Maintaining Plasticity via Continual Backprop' from 2021) https://scholar.google.com/scholar?q=Loss+of+Plasticity+in+Deep+Continual+Learning 3. On Warm-Starting Neural Network Training — Jordan T. Ash, Ryan P. Adams, 2020 (NeurIPS) https://scholar.google.com/scholar?q=On+Warm-Starting+Neural+Network+Training 4. The Primacy Bias in Deep Reinforcement Learning — Evgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon, Aaron Courville, 2022 (ICML) https://scholar.google.com/scholar?q=The+Primacy+Bias+in+Deep+Reinforcement+Learning 5. Understanding Plasticity in Neural Networks — Clare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Avila Pires, Razvan Pascanu, Will Dabney, 2023 (ICML) https://scholar.google.com/scholar?q=Understanding+Plasticity+in+Neural+Networks 6. RL's razor: Why online reinforcement learning forgets less — Idan Shenfeld, Jyothish Pari, Pulkit Agrawal, 2025 https://scholar.google.com/scholar?q=RL%27s+razor%3A+Why+online+reinforcement+learning+forgets+less 7. The Art of Scaling Reinforcement Learning Compute for LLMs (ScaleRL) — Devvrit Khatri, Lovish Madaan, Rishabh Tiwari, Rachit Bansal, Sai Surya Duvvuri, Manzil Zaheer, Inderjit S. Dhillon, David Brandfonbrener, Rishabh Agarwal, 2025 https://scholar.google.com/scholar?q=The+Art+of+Scaling+Reinforcement+Learning+Compute+for+LLMs+%28ScaleRL%29 8. Fine-tuning and prompt optimization: Two great steps that work better together (BetterTogether) — Dilara Soylu, Christopher Potts, Omar Khattab, 2024 https://scholar.google.com/scholar?q=Fine-tuning+and+prompt+optimization%3A+Two+great+steps+that+work+better+together+%28BetterTogether%29 9. Mitigating plasticity loss in continual reinforcement learning by reducing churn — Hongyao Tang, Johan Obando-Ceron, Pablo Samuel Castro, Aaron Courville, Glen Berseth, 2025 https://scholar.google.com/scholar?q=Mitigating+plasticity+loss+in+continual+reinforcement+learning+by+reducing+churn 10. What can you do when you have zero rewards during RL? — Jatin Prakash, Anirudh Buvanesh, 2025 https://scholar.google.com/scholar?q=What+can+you+do+when+you+have+zero+rewards+during+RL%3F Interactive Visualization: Learning, Fast and Slow: LLMs That Adapt Without Forgetting
-
679
Volatility Optimization Is Actually Bayesian Inference
This episode explores Kohei Honda's tutorial and survey "Model Predictive Control via Probabilistic Inference," which unifies two decades of scattered research—path integral control, reinforcement learning theory, and variational inference—into a single coherent framework called PI-MPC. The discussion traces why classical gradient- and Hessian-based MPC solvers break down on contact-rich robotics, learned neural dynamics, or discontinuous costs, and why the resulting fallback to naive random-shooting sampling collapses under the curse of dimensionality. The core argument is that reframing sampling-based MPC as inference over a distribution of good control sequences—rather than search for a single optimum—yields dramatic gains in sample efficiency and parallelizability, with MPPI's Boltzmann-weighted, temperature-controlled posterior serving as the paper's central worked example. Along the way, the hosts debate whether "inference" is meaningfully different from optimization, tracing how entropy terms in algorithms like Soft Actor-Critic emerge naturally from the probabilistic framing rather than being added as an exploration hack. Listeners interested in robotics, control theory, or the mathematical bridges between classical control and modern probabilistic ML will find the episode's account of why this synthesis only became practical with GPU-scale parallel rollouts particularly compelling. Sources: 1. Model Predictive Control via Probabilistic Inference: A Tutorial and Survey — Kohei Honda, 2025 http://arxiv.org/abs/2511.08019v4 2. Constrained Model Predictive Control: Stability and Optimality — D. Q. Mayne, J. B. Rawlings, C. V. Rao, P. O. M. Scokaert, 2000 https://scholar.google.com/scholar?q=Constrained+Model+Predictive+Control%3A+Stability+and+Optimality 3. A Survey of Industrial Model Predictive Control Technology — S. Joe Qin, Thomas A. Badgwell, 2003 https://scholar.google.com/scholar?q=A+Survey+of+Industrial+Model+Predictive+Control+Technology 4. Model Predictive Control: Theory and Practice — A Survey — Carlos E. Garcia, David M. Prett, Manfred Morari, 1989 https://scholar.google.com/scholar?q=Model+Predictive+Control%3A+Theory+and+Practice+%E2%80%94+A+Survey 5. Model Predictive Path Integral Control using Covariance Variable Importance Sampling — Grady Williams, Andrew Aldrich, Evangelos A. Theodorou, 2015 https://scholar.google.com/scholar?q=Model+Predictive+Path+Integral+Control+using+Covariance+Variable+Importance+Sampling 6. Information Theoretic MPC for Model-Based Reinforcement Learning — Grady Williams, Nolan Wagener, Brian Goldfain, Paul Drews, James M. Rehg, Byron Boots, Evangelos A. Theodorou, 2017 https://scholar.google.com/scholar?q=Information+Theoretic+MPC+for+Model-Based+Reinforcement+Learning 7. Robust Sampling Based Model Predictive Control with Sparse Objective Information — Grady Williams, Brian Goldfain, Paul Drews, Kamil Saigol, James M. Rehg, Evangelos A. Theodorou, 2018 https://scholar.google.com/scholar?q=Robust+Sampling+Based+Model+Predictive+Control+with+Sparse+Objective+Information 8. Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review — Sergey Levine, 2018 https://scholar.google.com/scholar?q=Reinforcement+Learning+and+Control+as+Probabilistic+Inference%3A+Tutorial+and+Review 9. Robot Trajectory Optimization using Approximate Inference — Marc Toussaint, 2009 https://scholar.google.com/scholar?q=Robot+Trajectory+Optimization+using+Approximate+Inference 10. Optimal Control as a Graphical Model Inference Problem — Hilbert J. Kappen, Vicenç Gómez, Manfred Opper, 2012 https://scholar.google.com/scholar?q=Optimal+Control+as+a+Graphical+Model+Inference+Problem 11. Variational Inference: A Review for Statisticians — David M. Blei, Alp Kucukelbir, Jon D. McAuliffe, 2017 https://scholar.google.com/scholar?q=Variational+Inference%3A+A+Review+for+Statisticians 12. Auto-Encoding Variational Bayes — Diederik P. Kingma, Max Welling, 2013 https://scholar.google.com/scholar?q=Auto-Encoding+Variational+Bayes 13. An Introduction to Variational Methods for Graphical Models — Michael I. Jordan, Zoubin Ghahramani, Tommi S. Jaakkola, Lawrence K. Saul, 1999 https://scholar.google.com/scholar?q=An+Introduction+to+Variational+Methods+for+Graphical+Models 14. Predictive Sampling: Real-Time Behaviour Synthesis with MuJoCo — Taylor Howell, Nimrod Gileadi, Saran Tunyasuvunakool, Kevin Zakka, Tom Erez, Yuval Tassa, 2022 https://scholar.google.com/scholar?q=Predictive+Sampling%3A+Real-Time+Behaviour+Synthesis+with+MuJoCo 15. STORM: An Integrated Framework for Fast Joint-Space Model-Predictive Control for Reactive Manipulation — Mohak Bhardwaj, Balakumar Sundaralingam, Arsalan Mousavian, Nathan D. Ratliff, Dieter Fox, Fabio Ramos, Byron Boots, 2021 https://scholar.google.com/scholar?q=STORM%3A+An+Integrated+Framework+for+Fast+Joint-Space+Model-Predictive+Control+for+Reactive+Manipulation 16. Information-Theoretic Model Predictive Control: Theory and Applications to Autonomous Driving — Grady Williams, Paul Drews, Brian Goldfain, James M. Rehg, Evangelos A. Theodorou, 2018 https://scholar.google.com/scholar?q=Information-Theoretic+Model+Predictive+Control%3A+Theory+and+Applications+to+Autonomous+Driving 17. Model-Based Diffusion for Trajectory Optimization — Chaoyi Pan, Zeji Yi, Guanya Shi, Guannan Qu, 2024 https://scholar.google.com/scholar?q=Model-Based+Diffusion+for+Trajectory+Optimization 18. TD-MPC2: Scalable, Robust World Models for Continuous Control — Nicklas Hansen, Hao Su, Xiaolong Wang, 2023 https://scholar.google.com/scholar?q=TD-MPC2%3A+Scalable%2C+Robust+World+Models+for+Continuous+Control 19. Recent Advances in Path Integral Control for Trajectory Optimization: An Overview in Theoretical and Algorithmic Perspectives — Muhammad Kazim, Jungee Hong, Min-Gyeom Kim, Kwang-Ki K. Kim, 2024 https://scholar.google.com/scholar?q=Recent+Advances+in+Path+Integral+Control+for+Trajectory+Optimization%3A+An+Overview+in+Theoretical+and+Algorithmic+Perspectives Interactive Visualization: Volatility Optimization Is Actually Bayesian Inference
-
678
TwinQuant: Manifold-Constrained Low-Rank Decomposition for 4-Bit Quantization
This episode explores TwinQuant, a 4-bit post-training quantization method for large language models that challenges a core assumption behind prior techniques like SVDQuant: that a weight matrix's important information can be captured in a small, fixed set of directions. The hosts explain how LLM weight outliers turn out to be spread across hundreds of directions rather than concentrated, forcing earlier low-rank decomposition approaches into an unwinnable tradeoff between speed and accuracy. They unpack TwinQuant's solution — learning the low-rank split itself via manifold optimization, using a true orthogonal (Stiefel manifold) rotation that folds cleanly into RMSNorm layers alongside a more flexible invertible (general linear) transform for layer-specific residual handling — plus a fused kernel designed to keep the approach fast at inference. Along the way, the conversation walks through foundational quantization vocabulary (PTQ, WxAy notation, mixed-precision splits) for listeners newer to the topic. It's a compelling listen for anyone tracking how far LLMs can be compressed without sacrificing accuracy, and why the math behind "which parts of a weight matrix matter" is more complicated than earlier compression work assumed. Sources: 1. TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization — Haodong Wang, Junjie Liu, Zicong Hong, Qianli Liu, Jian Lin, Song Guo, Xu Chen, 2026 http://arxiv.org/abs/2606.01556 2. Optimization Algorithms on Matrix Manifolds — P.-A. Absil, R. Mahony, R. Sepulchre, 2008 https://scholar.google.com/scholar?q=Optimization+Algorithms+on+Matrix+Manifolds 3. SpinQuant: LLM Quantization with Learned Rotations — Zechun Liu, Changsheng Zhao, Igor Fedorov, et al. (Meta AI), 2024 https://scholar.google.com/scholar?q=SpinQuant%3A+LLM+Quantization+with+Learned+Rotations 4. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs — Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, et al., 2024 https://scholar.google.com/scholar?q=QuaRot%3A+Outlier-Free+4-Bit+Inference+in+Rotated+LLMs 5. Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley Transform — Jun Li, Fuxin Li, Sinisa Todorovic, 2020 https://scholar.google.com/scholar?q=Efficient+Riemannian+Optimization+on+the+Stiefel+Manifold+via+the+Cayley+Transform 6. SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models — Li, M., Lin, Y., Zhang, Z., Cai, T., Li, X., Guo, J., Xie, E., Meng, C., Zhu, J.-Y., Han, S., 2025 https://scholar.google.com/scholar?q=SVDQuant%3A+Absorbing+Outliers+by+Low-Rank+Components+for+4-Bit+Diffusion+Models 7. FlatQuant: Flatness Matters for LLM Quantization — Sun, Y., Liu, R., Bai, H., Bao, H., Zhao, K., Li, Y., Yu, X., Hou, L., Yuan, C., Jiang, X., et al., 2025 https://scholar.google.com/scholar?q=FlatQuant%3A+Flatness+Matters+for+LLM+Quantization 8. OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models — Shao, W., Chen, M., Zhang, Z., Xu, P., Zhao, L., Li, Z., Zhang, K., Gao, P., Qiao, Y., Luo, P., 2024 https://scholar.google.com/scholar?q=OmniQuant%3A+Omnidirectionally+Calibrated+Quantization+for+Large+Language+Models Interactive Visualization: TwinQuant: Manifold-Constrained Low-Rank Decomposition for 4-Bit Quantization
-
677
TFGN: Replay-Free, Task-Free Continual Pre-Training at Scale
This episode explores TFGN, an architectural approach to continual pre-training of large language models that claims to solve catastrophic forgetting without four common crutches: replay buffers, task identifiers, small-scale toy benchmarks, and external penalty terms like Fisher-information regularization. The hosts trace the lineage of the forgetting problem back to 1989, explain why popular fixes like LoRA-based parameter-efficient fine-tuning don't actually address forgetting (they just shrink the blast radius), and why classic regularization methods like Elastic Weight Consolidation break down at billion-parameter scale. They also clarify why long-context windows and prompt-based knowledge aren't a substitute for genuinely updating model weights on massive, unbounded corpora like full codebases or legal archives. The conversation lays out TFGN's core mechanism as a dense, input-conditioned overlay operating inside each transformer block, contrasting it with sparse mixture-of-experts routing, and sets up backward transfer as the key metric for measuring whether old knowledge survives new training. Listeners interested in how production LLMs might eventually absorb new domains without expensive retraining or fragile adapter stacking will find the framing of this open problem sharply drawn. Sources: 1. TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale — Anurup Ganguli, 2026 http://arxiv.org/abs/2605.15053 2. Overcoming catastrophic forgetting in neural networks (EWC) — J. Kirkpatrick et al., 2017 https://scholar.google.com/scholar?q=Overcoming+catastrophic+forgetting+in+neural+networks+%28EWC%29 3. Loss of plasticity in deep continual learning — S. Dohare et al., 2024, Nature https://scholar.google.com/scholar?q=Loss+of+plasticity+in+deep+continual+learning 4. Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-Tuning — O. Y. L. Imanov, 2026, arXiv:2601.18699 https://scholar.google.com/scholar?q=Mechanistic+Analysis+of+Catastrophic+Forgetting+in+Large+Language+Models+During+Continual+Fine-Tuning 5. Examining Forgetting in Continual Pre-training of Aligned Large Language Models — C.-A. Li and H.-Y. Lee, 2024, arXiv:2401.03129 https://scholar.google.com/scholar?q=Examining+Forgetting+in+Continual+Pre-training+of+Aligned+Large+Language+Models 6. Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models — I. Abbes, G. Subbaraj, M. Riemer, et al., 2025, arXiv:2508.01908 https://scholar.google.com/scholar?q=Revisiting+Replay+and+Gradient+Alignment+for+Continual+Pre-Training+of+Large+Language+Models 7. Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thought — Y. Zhang, B. Tang, T. Ju, S. Duan, G. Liu, 2025, arXiv:2512.21711 https://scholar.google.com/scholar?q=Do+Latent+Tokens+Think%3F+A+Causal+and+Adversarial+Analysis+of+Chain-of-Continuous-Thought Interactive Visualization: TFGN: Replay-Free, Task-Free Continual Pre-Training at Scale
-
676
SnapStream: Taming KV-Cache Eviction for Dataflow Accelerators
This episode explores SnapStream, a technique from SambaNova Systems for compressing KV caches during long-sequence LLM decoding on dataflow accelerators, demonstrated at production scale with a 671-billion-parameter DeepSeek-R1 deployment running 128K-token context at over 1,800 tokens per second. The discussion covers why established training-free KV cache eviction methods like SnapKV and StreamingLLM have struggled to reach real deployments despite promising accuracy results: continuous batching makes it unclear when to trigger compression across requests at different lifecycle stages, and static-graph compilers used by dataflow accelerators can't easily accommodate the dynamic, variable-shaped operations that standard compression implementations rely on. It explains how SnapStream fuses SnapKV's attention-based token selection with StreamingLLM's sink-plus-sliding-window approach into a single fixed-size cache, splitting sequences into sink tokens, recent tokens, and a compressed middle section during prefill. The conversation is grounded in fundamentals—clarifying the prefill/decode split, why decode is memory-bound, and what makes dataflow accelerators architecturally different from GPUs—making it accessible to listeners unfamiliar with KV cache mechanics while still delivering a specific, hardware-grounded engineering story rather than a purely algorithmic one. Sources: 1. SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators — Jonathan Li, Nasim Farahini, Evgenii Iuliugin, Magnus Vesterlund, Christian Häggström, Guangtao Wang, Shubhangi Upasani, Ayush Sachdeva, Rui Li, Faline Fu, Chen Wu, Ayesha Siddiqua, John Long, Tuowen Zhao, Matheen Musaddiq, Håkan Zeffer, Yun Du, Mingran Wang, Qinghua Li, Bo Li, Urmish Thakker, Raghu Prabhakar, 2025 http://arxiv.org/abs/2511.03092 2. Plasticine: A Reconfigurable Architecture For Parallel Patterns — Raghu Prabhakar, Yaqi Zhang, David Koeplinger, Matt Feldman, Tian Zhao, Stefan Hadjis, Ardavan Pedram, Christos Kozyrakis, Kunle Olukotun, 2017 https://scholar.google.com/scholar?q=Plasticine%3A+A+Reconfigurable+Architecture+For+Parallel+Patterns 3. SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts — Raghu Prabhakar and SambaNova Systems architecture team, 2024 https://scholar.google.com/scholar?q=SN40L%3A+Scaling+the+AI+Memory+Wall+with+Dataflow+and+Composition+of+Experts 4. Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads — Dennis Abts, Jonathan Ross, Jonathan Sparling, Mark Wong-VanHaren, et al. (Groq), 2020 https://scholar.google.com/scholar?q=Think+Fast%3A+A+Tensor+Streaming+Processor+%28TSP%29+for+Accelerating+Deep+Learning+Workloads 5. In-Datacenter Performance Analysis of a Tensor Processing Unit — Norman P. Jouppi et al. (Google), 2017 https://scholar.google.com/scholar?q=In-Datacenter+Performance+Analysis+of+a+Tensor+Processing+Unit 6. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized LLM Serving — Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, H. Zhang, 2024 https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+LLM+Serving 7. Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference — J. Tang, Y. Zhao, K. Zhu, G. Xiao, B. Kasikci, S. Han, 2024 https://scholar.google.com/scholar?q=Quest%3A+Query-Aware+Sparsity+for+Efficient+Long-Context+LLM+Inference 8. InfLLM: Training-Free Long-Context Extrapolation with an Efficient Context Memory — C. Xiao, P. Zhang, X. Han, G. Xiao, Y. Lin, Z. Zhang, Z. Liu, M. Sun, 2024 https://scholar.google.com/scholar?q=InfLLM%3A+Training-Free+Long-Context+Extrapolation+with+an+Efficient+Context+Memory 9. DeepSeek-V3.2-Exp: Boosting Long-Context Efficiency with DeepSeek Sparse Attention — DeepSeek-AI, 2025 https://scholar.google.com/scholar?q=DeepSeek-V3.2-Exp%3A+Boosting+Long-Context+Efficiency+with+DeepSeek+Sparse+Attention 10. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study — S. Eyuboglu, R. Ehrlich, S. Arora, N. Guha, D. Zinsley, E. Liu, W. Tennien, A. Rudra, J. Zou, A. Mirhoseini, C. Re, 2025 https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+General-Purpose+Long+Context+Representations+via+Self-Study 11. LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction (SAGE-KV) — G. Wang, S. Upasani, C. Wu, D. Gandhi, J. Li, C. Hu, B. Li, U. Thakker, 2025 https://scholar.google.com/scholar?q=LLMs+Know+What+to+Drop%3A+Self-Attention+Guided+KV+Cache+Eviction+%28SAGE-KV%29 Interactive Visualization: SnapStream: Taming KV-Cache Eviction for Dataflow Accelerators
-
675
StrataCL: Fabric-Native Zero-Copy Comms for AI Supernodes
This episode explores StrataCL, a fabric-native communication library from researchers at Peking University, ICT-CAS, UCAS, Shanghai Jiao Tong University, and Huawei, tested on Huawei's CloudMatrix384 supernode. The discussion centers on how communication overhead — which the paper puts at 30-45% of end-to-end time in distributed LLM training and up to 50% at scale — can be cut by giving collectives and MoE routing true zero-copy access to application buffers on unified-memory fabrics, without breaking compatibility with frameworks like PyTorch and SGLang. A key insight is why buffer-centric libraries like NCCL and HCCL fall short even on fast unified-address fabrics, and how MoE dispatch/combine traffic exposes the limits of naive redesigns. The core technical contribution is registration-on-allocation: exploiting the multi-second gap between physical memory allocation and first use by a communication operator to move registration off the critical path entirely, asynchronously, the moment memory is mapped. The result is a 1.4x iteration-time speedup on a 512-die production training run with no changes to the model, optimizer, or data — pure systems engineering payoff. Sources: 1. StrataCL: Fabric-Native Communication Library for Production Supernodes — Tiancheng Hu, Jin Qin, Yuzheng Wang, Ke Liu, TangShengsheng Li, Sheng Wang, Zhongzhe Hu, Tianlun Hu, Wei Wang, Lijun Li, Jingbin Zhou, Xiaoming Bao, Hongwei Sun, Jieru Zhao, Huimin Cui, Tao Xie, Chenxi Wang, 2026 http://arxiv.org/abs/2607.26444 2. U-Net: A User-Level Network Interface for Parallel and Distributed Computing — Thorsten von Eicken, Anindya Basu, Vineet Buch, Werner Vogels, 1995 https://scholar.google.com/scholar?q=U-Net%3A+A+User-Level+Network+Interface+for+Parallel+and+Distributed+Computing 3. Design Guidelines for High Performance RDMA Systems — Anuj Kalia, Michael Kaminsky, David G. Andersen, 2016 https://scholar.google.com/scholar?q=Design+Guidelines+for+High+Performance+RDMA+Systems 4. Blink: Fast and Generic Collectives for Distributed ML — Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee, Jorgen Thelin, Nikhil Devanur, Ion Stoica, 2020 https://scholar.google.com/scholar?q=Blink%3A+Fast+and+Generic+Collectives+for+Distributed+ML 5. NVSHMEM (GPU-initiated, PGAS-style one-sided communication library) — NVIDIA (library/runtime, not a single academic paper), 2016 (initial release, iterated since) https://scholar.google.com/scholar?q=NVSHMEM+%28GPU-initiated%2C+PGAS-style+one-sided+communication+library%29 6. Collective Communication for 100k+ GPUs — Min Si, Pavan Balaji, Yongzhou Chen, et al., 2025 https://scholar.google.com/scholar?q=Collective+Communication+for+100k%2B+GPUs 7. SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA Offloading — Xingyi Li, Yadong Liu, Xiaojie Huang, et al., 2026 (NSDI 26) https://scholar.google.com/scholar?q=SwiftEP%3A+Accelerating+MoE+Inference+with+Buffer+Fusion+and+TMA+Offloading 8. PyTorch Symmetric Memory / NVSHMEM-style same-VA mirrored buffers — PyTorch Team / NVIDIA (NVSHMEM), 2024-2025 https://scholar.google.com/scholar?q=PyTorch+Symmetric+Memory+%2F+NVSHMEM-style+same-VA+mirrored+buffers 9. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al., 2023 (SOSP 23) https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention Interactive Visualization: StrataCL: Fabric-Native Zero-Copy Comms for AI Supernodes
-
674
MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA
This episode explores a challenge to conventional wisdom in parameter-efficient fine-tuning, examining a method called MiCA that inverts the logic behind LoRA (Low-Rank Adaptation). Rather than letting trainable weight-update matrices drift freely, as standard LoRA does, MiCA deliberately anchors one matrix to the minor singular-value directions of a weight matrix — the low-energy, rarely-used "corners" that classical compression theory says to discard — leaving those directions free for new knowledge rather than overwriting the dominant, pretrained-heavy subspace. The discussion traces the technique's lineage through SVD, the Eckart-Young-Mirsky theorem, PiSSA's SVD-based initialization, and Minor Component Analysis, framing MiCA's core bet: catastrophic forgetting during fine-tuning may stem from cramming new information into already-saturated high-energy directions. Listeners interested in the mechanics of efficient model adaptation, knowledge editing, and where the field's assumptions about "useless" weight-matrix structure might be wrong will find the debate over whether this is a genuine architectural insight or a narrower refinement of existing PEFT ideas especially engaging. Sources: 1. MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA https://arxiv.org/pdf/2604.01694 2. The Approximation of One Matrix by Another of Lower Rank — Carl Eckart, Gale Young, 1936 https://scholar.google.com/scholar?q=The+Approximation+of+One+Matrix+by+Another+of+Lower+Rank 3. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, 2021 https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models 4. PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models — Fanxu Meng, Zhaohui Wang, Muhan Zhang, 2024 https://scholar.google.com/scholar?q=PiSSA%3A+Principal+Singular+Values+and+Singular+Vectors+Adaptation+of+Large+Language+Models 5. AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning — Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, Tuo Zhao, 2023 https://scholar.google.com/scholar?q=AdaLoRA%3A+Adaptive+Budget+Allocation+for+Parameter-Efficient+Fine-Tuning 6. SOMA: Singular Value Decomposed Minor Components Adaptation for Domain Generalizable Representation Learning — Seokju Yun, Seunghye Chae, Dongheon Lee, Youngmin Ro, 2025 https://scholar.google.com/scholar?q=SOMA%3A+Singular+Value+Decomposed+Minor+Components+Adaptation+for+Domain+Generalizable+Representation+Learning 7. Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-Tuning — Yu-Ang Lee, Ching-Yun Ko, Pin-Yu Chen, Mi-Yen Yeh, 2026 https://scholar.google.com/scholar?q=Learning+Rate+Matters%3A+Vanilla+LoRA+May+Suffice+for+LLM+Fine-Tuning 8. DoRA: Weight-Decomposed Low-Rank Adaptation — Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, Min-Hung Chen, 2024 https://scholar.google.com/scholar?q=DoRA%3A+Weight-Decomposed+Low-Rank+Adaptation 9. Locating and Editing Factual Associations in GPT (ROME) — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022 https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT+%28ROME%29 10. Editing Models with Task Arithmetic — Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, Ali Farhadi, 2023 https://scholar.google.com/scholar?q=Editing+Models+with+Task+Arithmetic Interactive Visualization: MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA
-
673
Move the Query, Not the Cache: MLA Rewrites GPU Fabric Attention Routing
This episode explores cross-instance attention in disaggregated LLM serving, focusing on the surprising size inversion created by Multi-head Latent Attention: a routed decoding query shrinks to roughly a kilobyte while the cache chunk it must read can balloon to 61 megabytes across layers, upending the old assumption that query and cache are comparably sized. The discussion traces why this scenario is becoming routine — providers sharing precomputed caches for large corpora that outgrow a single GPU's memory, and agentic workloads where many sub-agents query one oversized shared prefix — and lays out the three possible strategies (route, fetch, or recompute locally) for handling the mismatch, including how sparse indexers further shrink the routable unit to scattered top-k blocks. A key thread examines device-initiated RDMA via IBGDA, challenging the intuition that skipping the CPU proxy is automatically faster: prior work on tiny mixture-of-experts messages actually found IBGDA slower, but the paper's controlled test on kilobyte-scale attention traffic shows the CPU-proxy path is 40% slower at the median and over 50% slower at steady state. Listeners interested in GPU networking, KV-cache architecture, or the practical plumbing behind large-scale LLM inference will find the paper's empirical resolution of a previously untested assumption particularly compelling. Sources: 1. Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics — Bole Ma, Jan Eitzinger, Harald Köstler, Gerhard Wellein, 2026 http://arxiv.org/abs/2606.01502 2. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI (research team), 2024 https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model 3. DeepSeek-V3 Technical Report — DeepSeek-AI (research team), 2024 https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report 4. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, Sumit Sanghai, 2023 https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints 5. TransMLA: Multi-Head Latent Attention Is All You Need — Fanxu Meng, Zengwei Yao, Muhan Zhang, et al., 2025 https://scholar.google.com/scholar?q=TransMLA%3A+Multi-Head+Latent+Attention+Is+All+You+Need 6. Improving Network Performance of HPC Systems Using NVIDIA Magnum IO NVSHMEM and GPUDirect Async — NVIDIA (NVSHMEM / Magnum IO engineering team), 2023 https://scholar.google.com/scholar?q=Improving+Network+Performance+of+HPC+Systems+Using+NVIDIA+Magnum+IO+NVSHMEM+and+GPUDirect+Async 7. Efficient Inter-node MPI Communication using GPUDirect RDMA for InfiniBand Clusters with NVIDIA GPUs — Sreeram Potluri, Khaled Hamidouche, Akshay Venkatesh, Devendar Bureddy, Dhabaleswar K. Panda, 2013 https://scholar.google.com/scholar?q=Efficient+Inter-node+MPI+Communication+using+GPUDirect+RDMA+for+InfiniBand+Clusters+with+NVIDIA+GPUs 8. DeepEP: an efficient expert-parallel communication library (and related DeepSeek-V3 Technical Report communication sections) — DeepSeek-AI (research/infra team), 2025 / 2024 https://scholar.google.com/scholar?q=DeepEP%3A+an+efficient+expert-parallel+communication+library+%28and+related+DeepSeek-V3+Technical+Report+communication+sections%29 9. Introducing OpenSHMEM: SHMEM for the PGAS Community — Barbara Chapman, Tony Curtis, Swaroop Pophale, et al., 2010 https://scholar.google.com/scholar?q=Introducing+OpenSHMEM%3A+SHMEM+for+the+PGAS+Community 10. Preble: Efficient Distributed Prompt Scheduling for LLM Serving — Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, Yiying Zhang, 2024 https://scholar.google.com/scholar?q=Preble%3A+Efficient+Distributed+Prompt+Scheduling+for+LLM+Serving 11. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu, Heyang Huang, Junhao Hu, Jiang Xu, Xusheng Chen, Tao Xie, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, Yizhou Shan, 2024 https://scholar.google.com/scholar?q=MemServe%3A+Context+Caching+for+Disaggregated+LLM+Serving+with+Elastic+Memory+Pool 12. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, Beidi Chen, 2023 https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models Interactive Visualization: Move the Query, Not the Cache: MLA Rewrites GPU Fabric Attention Routing
-
672
Silicon Showdown: GPU vs Apple Silicon LLM Inference Limits
This episode examines "Silicon Showdown," a study comparing Nvidia discrete-GPU and Apple unified-memory architectures for running large language models on consumer hardware, tested across model sizes from 1.5 billion to 80 billion parameters. It explains why Nvidia's VRAM Wall forces a stark trade-off between quantizing models down or offloading to slower system RAM across a PCIe bottleneck, while Apple's unified memory pool lets large models load fully without that penalty, at the cost of slower per-byte bandwidth. The discussion breaks down the competing software stacks—Nvidia's TensorRT-LLM with its new NVFP4 format and split-backend behavior, Apple's compilation-free MLX, and the cross-platform GGUF fallback from llama.cpp—and how each shapes real-world performance on metrics like time-to-first-token and tokens per joule. The episode highlights a gap in existing benchmarks like MLPerf and vLLM research, which focus on data-center throughput rather than the moment a model outgrows a single consumer GPU's memory. Listeners interested in running frontier open-weight models like Llama-3.3-70B or Qwen3-Next-80B on their own hardware will find a grounded, hardware-specific account of where each platform's approach breaks down. Sources: 1. Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference — Abdurrahman Javat, Allan Kazakov, 2026 http://arxiv.org/abs/2605.00519 2. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers — Frantar, Ashkboos, Hoefler, Alistarh, 2022 https://scholar.google.com/scholar?q=GPTQ%3A+Accurate+Post-Training+Quantization+for+Generative+Pre-trained+Transformers 3. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — Lin, Tang, Tang, Yang, Dang, Han, 2023 https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration 4. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale — Dettmers, Lewis, Belkada, Zettlemoyer, 2022 https://scholar.google.com/scholar?q=LLM.int8%28%29%3A+8-bit+Matrix+Multiplication+for+Transformers+at+Scale 5. Mixtral of Experts — Jiang et al. (Mistral AI), 2024 https://scholar.google.com/scholar?q=Mixtral+of+Experts 6. Fast Inference from Transformers via Speculative Decoding — Leviathan, Kalman, Matias, 2023 https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding 7. Efficient Memory Management for Large Language Model Serving with PagedAttention — Kwon et al., 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention Interactive Visualization: Silicon Showdown: GPU vs Apple Silicon LLM Inference Limits
-
671
Model Predictive Control's Real-Time Structure, from Chapter to Cockpit
Sitting down with a two-author-plus-one chemical-engineering-rooted textbook on Model Predictive Control, this episode unpacks why the field treats real-time feasibility as a hard constraint rather than a nice-to-have — walking through how MPC re-solves an optimization problem from scratch every control cycle using a known dynamics model, with no learning or reward signal involved. The discussion centers on the structural trick that makes this tractable on embedded hardware: exploiting the block-banded, time-local coupling of the problem via Riccati recursion or condensing to cut a naive O(N³) solve down to O(N), and how Diehl, Bock, and Schlöder's 2005 real-time iteration scheme turned this from a lab curiosity into something a drone or engine controller can rerun dozens of times per second. It also covers moving horizon estimation as the optimization-based counterpart to the Kalman filter, explaining why MHE can enforce physical constraints a Kalman filter can't, and why the book cuts particle filtering from its main text once state dimensionality climbs past five. Listeners get a clear picture of why the same machinery underlies powered-descent guidance, automotive control, and legged robotics — not as a trend, but as the only approach that reliably meets millisecond-scale deadlines. Sources: 1. Model Predictive Control's Real-Time Structure, from Chapter to Cockpit https://sites.engineering.ucsb.edu/~jbraw/mpc/MPC-book-2nd-edition-4th-printing.pdf 2. A Real-Time Iteration Scheme for Nonlinear Optimization in Optimal Feedback Control — Moritz Diehl, Hans Georg Bock, Johannes P. Schlöder, 2005 https://scholar.google.com/scholar?q=A+Real-Time+Iteration+Scheme+for+Nonlinear+Optimization+in+Optimal+Feedback+Control 3. CasADi: A Software Framework for Nonlinear Optimization and Optimal Control — Joel A. E. Andersson, Joris Gillis, Greg Horn, James B. Rawlings, Moritz Diehl, 2019 https://scholar.google.com/scholar?q=CasADi%3A+A+Software+Framework+for+Nonlinear+Optimization+and+Optimal+Control 4. acados: A Modular Open-Source Framework for Fast Embedded Optimal Control — Robin Verschueren, Gianluca Frison, Dimitris Kouzoupis, Jonathan Frey, Niels van Duijkeren, Andrea Zanelli, Branimir Novoselnik, Thivaharan Albin, Rien Quirynen, Moritz Diehl, 2022 https://scholar.google.com/scholar?q=acados%3A+A+Modular+Open-Source+Framework+for+Fast+Embedded+Optimal+Control 5. On the Implementation of an Interior-Point Filter Line-Search Algorithm for Large-Scale Nonlinear Programming — Andreas Wächter, Lorenz T. Biegler, 2006 https://scholar.google.com/scholar?q=On+the+Implementation+of+an+Interior-Point+Filter+Line-Search+Algorithm+for+Large-Scale+Nonlinear+Programming 6. Constrained Linear State Estimation — A Moving Horizon Approach — Christopher V. Rao, James B. Rawlings, Jay H. Lee, 2001 https://scholar.google.com/scholar?q=Constrained+Linear+State+Estimation+%25E2%2580%2594+A+Moving+Horizon+Approach 7. Constrained State Estimation for Nonlinear Discrete-Time Systems: Stability and Moving Horizon Approximations — Christopher V. Rao, James B. Rawlings, David Q. Mayne, 2003 https://scholar.google.com/scholar?q=Constrained+State+Estimation+for+Nonlinear+Discrete-Time+Systems%3A+Stability+and+Moving+Horizon+Approximations 8. Moving-Horizon State Estimation for Nonlinear Discrete-Time Systems: New Stability Results and Approximation Schemes — Angelo Alessandri, Marco Baglietto, Giorgio Battistelli, 2008 https://scholar.google.com/scholar?q=Moving-Horizon+State+Estimation+for+Nonlinear+Discrete-Time+Systems%3A+New+Stability+Results+and+Approximation+Schemes 9. Stochastic Model Predictive Control: An Overview and Perspectives for Future Research — Ali Mesbah, 2016 https://scholar.google.com/scholar?q=Stochastic+Model+Predictive+Control%3A+An+Overview+and+Perspectives+for+Future+Research 10. Stochastic Linear Model Predictive Control with Chance Constraints — A Review — Marcello Farina, Luca Giulioni, Riccardo Scattolini, 2016 https://scholar.google.com/scholar?q=Stochastic+Linear+Model+Predictive+Control+with+Chance+Constraints+%25E2%2580%2594+A+Review 11. Learning-Based Model Predictive Control: Toward Safe Learning in Control — Lukas Hewing, Kim P. Wabersich, Marcel Menner, Melanie N. Zeilinger, 2020 https://scholar.google.com/scholar?q=Learning-Based+Model+Predictive+Control%3A+Toward+Safe+Learning+in+Control 12. Architectures for Distributed and Hierarchical Model Predictive Control — A Review — Riccardo Scattolini, 2009 https://scholar.google.com/scholar?q=Architectures+for+Distributed+and+Hierarchical+Model+Predictive+Control+%25E2%2580%2594+A+Review 13. Distributed MPC Strategies with Application to Power System Automatic Generation Control — Aswin N. Venkat, Ian A. Hiskens, James B. Rawlings, Stephen J. Wright, 2008 https://scholar.google.com/scholar?q=Distributed+MPC+Strategies+with+Application+to+Power+System+Automatic+Generation+Control 14. Distributed Model Predictive Control: A Tutorial Review and Future Research Directions — Panagiotis D. Christofides, Riccardo Scattolini, David Muñoz de la Peña, Jinfeng Liu, 2013 https://scholar.google.com/scholar?q=Distributed+Model+Predictive+Control%3A+A+Tutorial+Review+and+Future+Research+Directions 15. acados — a modular open-source framework for fast embedded optimal control — R. Verschueren, G. Frison, D. Kouzoupis, N. van Duijkeren, A. Zanelli, B. Novoselnik, T. Albin, R. Quirynen, M. Diehl, 2022 https://scholar.google.com/scholar?q=acados+%25E2%2580%2594+a+modular+open-source+framework+for+fast+embedded+optimal+control 16. The scenario approach to robust control design — G.C. Calafiore, M.C. Campi, 2006 https://scholar.google.com/scholar?q=The+scenario+approach+to+robust+control+design 17. Stability of nonstationary receding horizon control (foundational stability result for stochastic MPC) — D. Chatterjee, J. Lygeros, 2015 https://scholar.google.com/scholar?q=Stability+of+nonstationary+receding+horizon+control+%28foundational+stability+result+for+stochastic+MPC%29 18. Robust MPC and dissipativity-based analysis for stochastic constrained systems (source of Assumption 3.22, stochastic MPC Version 2) — D.Q. Mayne, P. Falugi, 2019 https://scholar.google.com/scholar?q=Robust+MPC+and+dissipativity-based+analysis+for+stochastic+constrained+systems+%28source+of+Assumption+3.22%2C+stochastic+MPC+Version+2%29 19. A model predictive control framework for industrial turbodiesel engine control (source system for the nonlinear distributed MPC example) — B.T. Stewart, A.N. Venkat, J.B. Rawlings, S.J. Wright, G. Pannocchia (2011 IEEE CDC paper referenced as Stewart et al. 2011), 2011 https://scholar.google.com/scholar?q=A+model+predictive+control+framework+for+industrial+turbodiesel+engine+control+%28source+system+for+the+nonlinear+distributed+MPC+example%29 Interactive Visualization: Model Predictive Control's Real-Time Structure, from Chapter to Cockpit
-
670
Making Every Verified Token Count in MoE Speculative Decoding
This episode explores adaptive verification for speculative decoding when the target model is a sparse Mixture-of-Experts (MoE) system rather than a dense transformer, focusing on the paper "Making Every Verified Token Count." The discussion traces the lineage from Leviathan et al.'s original speculative decoding through tree-based drafting methods like Medusa and EAGLE-3, then explains why MoE architectures break a core assumption: since different draft-tree branches can route to entirely different experts, verifying a tree means loading every expert any branch touched. Drawing on the paper's benchmarks across three MoE models (including Qwen3-30B-A3B), the hosts unpack the striking finding that verification alone consumes 79-89% of per-iteration decoding latency once trees grow past thirty nodes — flipping the "verification is nearly free" pitch that made speculative decoding attractive in the first place. Listeners interested in LLM inference serving, GPU memory-bandwidth bottlenecks, or the practical tradeoffs of deploying sparse MoE models will find the episode's breakdown of why dense-model intuition fails on MoE targets especially clarifying. Sources: 1. Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding — Lehan Pan, Ziyang Tao, Ruoyu Pang, Xiao Wang, Jianjun Zhao, Yanyong Zhang, 2026 http://arxiv.org/abs/2605.00342 2. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023 https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding 3. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads — Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, Tri Dao, 2024 https://scholar.google.com/scholar?q=Medusa%3A+Simple+LLM+Inference+Acceleration+Framework+with+Multiple+Decoding+Heads 4. EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2025 https://scholar.google.com/scholar?q=EAGLE-3%3A+Scaling+up+Inference+Acceleration+of+Large+Language+Models+via+Training-Time+Test 5. Mixtral of Experts — Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, and the Mistral AI team, 2024 https://scholar.google.com/scholar?q=Mixtral+of+Experts 6. Utility-driven speculative decoding for mixture-of-experts — Anish Saxena, Po-An Tsai, Hritvik Taneja, Aamer Jaleel, Moinuddin Qureshi, 2025 https://scholar.google.com/scholar?q=Utility-driven+speculative+decoding+for+mixture-of-experts 7. MoE-Spec: Expert budgeting for efficient speculative decoding — Bradley McDanel, Steven Li, Sruthikesh Surineni, Harshit Khaitan, 2026 https://scholar.google.com/scholar?q=MoE-Spec%3A+Expert+budgeting+for+efficient+speculative+decoding 8. ECHO: Elastic speculative decoding with sparse gating for high-concurrency scenarios — Xinyi Hu, Yuhao Shen, Baolin Zhang, Hengxin Zhang, Jun Dai, Shuang Ge, Lei Chen, Yue Li, Mingcheng Wan, 2026 https://scholar.google.com/scholar?q=ECHO%3A+Elastic+speculative+decoding+with+sparse+gating+for+high-concurrency+scenarios 9. MoESD: Unveil speculative decoding's potential for accelerating sparse MoE — Zongle Huang, Lei Zhu, Zongyuan Zhan, Ting Hu, Weikai Mao, Xianzhi Yu, Yongpan Liu, Tianyu Zhang, 2025 https://scholar.google.com/scholar?q=MoESD%3A+Unveil+speculative+decoding%27s+potential+for+accelerating+sparse+MoE Interactive Visualization: Making Every Verified Token Count in MoE Speculative Decoding
-
669
FreeAct: Rethinking One-to-One Transforms for LLM Quantization
This episode explores FreeAct, a new approach to quantizing large language models down to 4-bit weights and activations (W4A4), presented by researchers from the National University of Singapore, Huawei Technology, and Central South University. The discussion traces how prior methods like QuaRot and FlatQuant rely on a rigid one-to-one pairing between a rotation matrix applied to activations and its exact inverse applied to weights — an assumption that breaks down for diffusion language models, where masked and unmasked tokens have different statistical profiles, and for multimodal models mixing vision and text tokens through the same layers. The hosts unpack the outlier-channel problem that makes activation quantization so much harder than weight quantization, tracing it back to Dettmers' LLM.int8 findings, and explain how FreeAct exploits a linear-algebra insight — dubbed Proposition 1 — showing that rank-deficient activation matrices allow a whole family of transformations rather than a single exact inverse, enabling different token types to use different activation-side matrices while keeping one shared weight-side transform. It's a compelling listen for anyone tracking how quantization techniques are adapting to increasingly heterogeneous token streams in modern AI systems. Sources: 1. FreeAct: Freeing Activations for LLM Quantization — Xiaohao Liu, Xiaobo Xia, Manyi Zhang, Ji-Fu Li, Xianzhi Yu, Fei Shen, Xiu Su, See-Kiong Ng, Tat-Seng Chua, 2026 http://arxiv.org/abs/2603.01776 2. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale — Tim Dettmers, Mike Lewis, Younes Belkada, Luke Zettlemoyer, 2022 https://scholar.google.com/scholar?q=LLM.int8%28%29%3A+8-bit+Matrix+Multiplication+for+Transformers+at+Scale 3. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models — Guangxuan Xiao, Ji Lin, Mickael Seznec, Julien Demouth, Song Han, 2023 https://scholar.google.com/scholar?q=SmoothQuant%3A+Accurate+and+Efficient+Post-Training+Quantization+for+Large+Language+Models 4. Atom: Low-bit Quantization for Efficient and Accurate LLM Serving — Yilong Zhao, Chien-Yu Lin, Kan Zhu, et al., 2024 https://scholar.google.com/scholar?q=Atom%3A+Low-bit+Quantization+for+Efficient+and+Accurate+LLM+Serving 5. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs — Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian Croci, et al., 2024 https://scholar.google.com/scholar?q=QuaRot%3A+Outlier-Free+4-Bit+Inference+in+Rotated+LLMs 6. FlatQuant: Flatness Matters for LLM Quantization — Yuxuan Sun, et al., 2025 https://scholar.google.com/scholar?q=FlatQuant%3A+Flatness+Matters+for+LLM+Quantization 7. SpinQuant: LLM Quantization with Learned Rotations — Zechun Liu, Changsheng Zhao, et al. (Meta AI), 2024 https://scholar.google.com/scholar?q=SpinQuant%3A+LLM+Quantization+with+Learned+Rotations 8. QuIP: 2-Bit Quantization of Large Language Models With Guarantees — Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De Sa, 2023 https://scholar.google.com/scholar?q=QuIP%3A+2-Bit+Quantization+of+Large+Language+Models+With+Guarantees 9. MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Static Quantization — Jiangyong Yu, Sifan Zhou, Dawei Yang, et al., 2025 https://scholar.google.com/scholar?q=MQuant%3A+Unleashing+the+Inference+Potential+of+Multimodal+Large+Language+Models+via+Static+Quantization 10. DLLMQuant: Quantizing Diffusion-based Large Language Models — Chen Xu, Dan Yang, 2025 https://scholar.google.com/scholar?q=DLLMQuant%3A+Quantizing+Diffusion-based+Large+Language+Models 11. Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models — Tianao Zhang, Zhiteng Li, Xianglong Yan, Haotong Qin, Yong Guo, Yulun Zhang, 2025 https://scholar.google.com/scholar?q=Quant-dLLM%3A+Post-Training+Extreme+Low-Bit+Quantization+for+Diffusion+Large+Language+Models 12. DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs — Haokun Lin, Haobo Xu, Yichen Wu, et al., 2024 https://scholar.google.com/scholar?q=DuQuant%3A+Distributing+Outliers+via+Dual+Transformation+Makes+Stronger+Quantized+LLMs Interactive Visualization: FreeAct: Rethinking One-to-One Transforms for LLM Quantization
-
668
Global Memory Bloat in Long-Context LLM Serving
This episode surveys how large language model serving systems manage the key-value cache — the memory storing every token's key and value vectors — as it has grown from a disposable per-request tensor into a resource actively managed, moved, and contended for across GPUs, nodes, and storage tiers. Drawing on a Texas Tech University paper classifying over thirty existing systems, the hosts unpack the arithmetic behind why KV cache footprint balloons with long context windows (reaching roughly 40 gigabytes for a single 128K-token request on a 70-billion-parameter model) and why bandwidth, not just capacity, becomes the real bottleneck during decode. They trace the field's foundational shift back to PagedAttention, the vLLM technique that introduced OS-style paging for KV memory, and explain how nearly every later system builds on its block-table abstraction. The conversation then turns to a four-dimensional taxonomy — locality, lifetime, ownership, and transport — used to organize the design space, highlighting a striking gap where two of five lifetime categories contain zero real-world systems. Listeners interested in LLM infrastructure, memory hierarchies, or the practical limits of long-context and agentic serving will find a clear framework for reasoning about a problem that's easy to underestimate with a single "the cache grows" intuition. Sources: 1. From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving — Jie Li, Tongyang Wang, Yong Chen, 2026 http://arxiv.org/abs/2607.02574 2. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI, 2024 https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model 3. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon et al., 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 4. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin et al., 2024 https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving 5. Medusa / EAGLE speculative decoding work (Cai et al. 2024; Li et al. 2024) — Tianle Cai et al.; Yuhui Li et al., 2024 https://scholar.google.com/scholar?q=Medusa+%2F+EAGLE+speculative+decoding+work+%28Cai+et+al.+2024%3B+Li+et+al.+2024%29 6. S-LoRA: Serving Thousands of Concurrent LoRA Adapters — Ying Sheng et al., 2023/2024 https://scholar.google.com/scholar?q=S-LoRA%3A+Serving+Thousands+of+Concurrent+LoRA+Adapters Interactive Visualization: Global Memory Bloat in Long-Context LLM Serving
-
667
DualDecoder: Fixing GPU Memory Bloat in Long-Context KV Cache Offloading
This episode covers "DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch," which tackles a hidden cost in dynamic sparse KV-cache systems: the GPU-resident bookkeeping state (landmarks, reconstructed keys) used to make host-memory offloading fast can itself consume up to 64% of GPU memory — 8.5 times larger than the actual sparse KV entries it's meant to retrieve. Drawing on the lineage from H2O's heavy-hitter observation to ShadowKV's landmark-based retrieval, the discussion explains how this auxiliary overhead quietly erodes the memory savings these systems promise, with ShadowKV reaching only 6.7% of its idealized batch-size capacity on a 32-billion-parameter model. DualDecoder's proposed fix is predictive prefetching: rather than permanently parking retrieval-support state on the GPU, it predicts the next decoding step's needs one step ahead and pulls entries from host memory just in time. Listeners interested in LLM inference efficiency will find a concrete, measured account of how a fix for one memory wall can quietly build a smaller one right next to it — and a proposed way out. Sources: 1. DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch — Zuning Liang, Zhiyi Yao, Qi Chen, Yuedong Xu, Hao Dai, Zhiqiang Ding, Tongkai Yang, Jinlong Hou, Yuan Cheng, 2026 http://arxiv.org/abs/2607.26475 2. SpeContext: Enabling Efficient Long-context Reasoning with Speculative Context Sparsity in LLMs — Jiaming Xu, Jiayi Pan, Hanzhen Wang, Yongkang Zhou, Jiancai Ye, Yu Wang, Guohao Dai, 2026 https://scholar.google.com/scholar?q=SpeContext%3A+Enabling+Efficient+Long-context+Reasoning+with+Speculative+Context+Sparsity+in+LLMs 3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, et al., 2023 https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models 4. RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval — Di Liu, Meng Chen, Baotong Lu, et al., 2024 https://scholar.google.com/scholar?q=RetrievalAttention%3A+Accelerating+Long-Context+LLM+Inference+via+Vector+Retrieval 5. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu, Heyang Huang, Junhao Hu, et al., 2024 https://scholar.google.com/scholar?q=MemServe%3A+Context+Caching+for+Disaggregated+LLM+Serving+with+Elastic+Memory+Pool 6. Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM) — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al., 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention+%28vLLM%29 Interactive Visualization: DualDecoder: Fixing GPU Memory Bloat in Long-Context KV Cache Offloading
-
666
Data Temporality's Hidden Impact on LLM Pretraining
This episode explores why open-weight LLMs like Llama 3.1, Gemma3, Qwen3, and Olmo3 systematically lose 11–39% relative accuracy on facts from 2023–2024 compared to facts from 2020–2021, even though the more recent data falls within their training window. Drawing on Kyutai's paper "Understanding Data Temporality Impact on Large Language Models Pre-training," the discussion traces this "knowledge horizon gap" to a design choice baked into standard pretraining: corpora from many years are pooled and globally shuffled before training, erasing any timestamp signal and letting older, more frequently re-crawled data dominate. The hosts connect this to learning-rate decay schedules, arguing that data seen late in training — when updates are small and durable — gets imprinted far more strongly than data seen early, so chronological ordering (feeding snapshots 2018 through 2025 in sequence) could exploit that same mechanism to anchor recent facts instead of losing them. They situate the work against Zhao et al.'s "Set the Clock" research and Bengio's foundational curriculum-learning ideas, framing chronological training as a strikingly cheap intervention — same tokens, same compute, same architecture — for a problem the field has largely ignored. It's a compelling listen for anyone puzzling over why "knowledge cutoff" claims don't match what models actually seem to know. Sources: 1. Understanding Data Temporality Impact on Large Language Models Pre-training — Hippolyte Pilchen, Romain Fabre, Franck Signe Talla, Patrick Perez, Edouard Grave, 2026 http://arxiv.org/abs/2605.22769 2. Set the Clock: Temporal Alignment of Pretrained Language Models — Bowen Zhao, Zander Brumbaugh, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, 2024 https://scholar.google.com/scholar?q=Set+the+Clock%3A+Temporal+Alignment+of+Pretrained+Language+Models 3. Time-Aware Language Models as Temporal Knowledge Bases — Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, William W. Cohen, 2022 https://scholar.google.com/scholar?q=Time-Aware+Language+Models+as+Temporal+Knowledge+Bases 4. TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language Models — Joel Jang, Seonghyeon Ye, Changho Lee, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Minjoon Seo, 2022 https://scholar.google.com/scholar?q=TemporalWiki%3A+A+Lifelong+Benchmark+for+Training+and+Evaluating+Ever-Evolving+Language+Models 5. RealTime QA: What's the Answer Right Now? — Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Yutaro Yamada, Deqing Fu, Tushar Khot, Ashish Sabharwal, Rik Koncel-Kedziorski, Yejin Choi, Noah A. Smith, Kentaro Inui, 2022 https://scholar.google.com/scholar?q=RealTime+QA%3A+What%27s+the+Answer+Right+Now%3F 6. Curriculum Learning — Yoshua Bengio, Jérôme Louradour, Ronan Collobert, Jason Weston, 2009 https://scholar.google.com/scholar?q=Curriculum+Learning 7. TimeLMs: Diachronic Language Models from Twitter — Daniel Loureiro, Francesco Barbieri, Leonardo Neves, Luis Espinosa Anke, Jose Camacho-Collados, 2022 https://scholar.google.com/scholar?q=TimeLMs%3A+Diachronic+Language+Models+from+Twitter 8. In-Context Pretraining: Language Modeling Beyond Document Boundaries — Weijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou, Margaret Li, Xi Victoria Lin, Noah A. Smith, Luke Zettlemoyer, Wen-tau Yih, Mike Lewis, 2023 https://scholar.google.com/scholar?q=In-Context+Pretraining%3A+Language+Modeling+Beyond+Document+Boundaries 9. Towards Continual Knowledge Learning of Language Models — Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, Minjoon Seo, 2022 https://scholar.google.com/scholar?q=Towards+Continual+Knowledge+Learning+of+Language+Models 10. TiC-LM: A web-scale benchmark for time-continual LLM pretraining — Li, J., Armandpour, M., Mirzadeh, I., Mehta, S., Shankar, V., Vemulapalli, R., Bengio, S., Tuzel, O., Farajtabar, M., Pouransari, H., Faghri, F., 2025 https://scholar.google.com/scholar?q=TiC-LM%3A+A+web-scale+benchmark+for+time-continual+LLM+pretraining 11. How do language models learn facts? Dynamics, curricula and hallucinations — Zucchet, N., Bornschein, J., Chan, S. C., Lampinen, A. K., Pascanu, R., De, S., 2025 https://scholar.google.com/scholar?q=How+do+language+models+learn+facts%3F+Dynamics%2C+curricula+and+hallucinations 12. Data mixing can induce phase transitions in knowledge acquisition — Gu, X., Lyu, K., Li, J., Zhang, J., 2026 https://scholar.google.com/scholar?q=Data+mixing+can+induce+phase+transitions+in+knowledge+acquisition 13. TiMoE: Time-aware mixture of language experts — Faro, R., Fan, D., Alphaidze, T., Jaggi, M., 2025 https://scholar.google.com/scholar?q=TiMoE%3A+Time-aware+mixture+of+language+experts 14. Does your data spark joy? Performance gains from domain upsampling at the end of training — Blakeney, C., Paul, M., Larsen, B. W., Owen, S., Frankle, J., 2024 https://scholar.google.com/scholar?q=Does+your+data+spark+joy%3F+Performance+gains+from+domain+upsampling+at+the+end+of+training Interactive Visualization: Data Temporality's Hidden Impact on LLM Pretraining
-
665
Adapting Without Forgetting: A Lifelong Learning Roadmap for LLM Agents
This episode explores "Lifelong Learning of Large Language Model based Agents: A Roadmap," a survey examining how AI agents can continuously adapt to changing environments without losing prior knowledge. The discussion centers on the stability-plasticity dilemma—the tension between preserving learned capabilities and remaining flexible enough to absorb new information—and how this classical problem from connectionist neuroscience resurfaces in a new form for modern agents that rarely fine-tune their underlying weights. Key arguments include the concept of "functional forgetting," where information technically persists in vector stores but becomes practically inaccessible if retrieval or context limits fail to surface it, and a four-part memory taxonomy spanning working, episodic, semantic, and parametric memory. The hosts also trace how this survey synthesizes and extends two separate research lineages—internal-knowledge-focused LLM surveys and agent-architecture surveys—into a unified framework modeled as a goal-conditioned POMDP. Listeners interested in why coding assistants, web-browsing agents, and other AI tools degrade over time as their environments shift will find concrete framing for that problem here. Sources: 1. Lifelong Learning of Large Language Model based Agents: A Roadmap — Junhao Zheng, Chengming Shi, Xidi Cai, Qiuke Li, Duzhen Zhang, Chenxing Li, Dong Yu, Qianli Ma, 2025 http://arxiv.org/abs/2501.07278 2. Overcoming Catastrophic Forgetting in Neural Networks — James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, et al. (DeepMind), 2017 https://scholar.google.com/scholar?q=Overcoming+Catastrophic+Forgetting+in+Neural+Networks 3. Continual Lifelong Learning with Neural Networks: A Review — German I. Parisi, Ronald Kemker, Jose L. Part, Christopher Kanan, Stefan Wermter, 2019 https://scholar.google.com/scholar?q=Continual+Lifelong+Learning+with+Neural+Networks%3A+A+Review 4. Generative Agents: Interactive Simulacra of Human Behavior — Joon Sung Park, Joseph O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein (Stanford / Google), 2023 https://scholar.google.com/scholar?q=Generative+Agents%3A+Interactive+Simulacra+of+Human+Behavior 5. Voyager: An Open-Ended Embodied Agent with Large Language Models — Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, Anima Anandkumar (NVIDIA, Caltech, UT Austin), 2023 https://scholar.google.com/scholar?q=Voyager%3A+An+Open-Ended+Embodied+Agent+with+Large+Language+Models 6. Towards Lifelong Learning of Large Language Models: A Survey — J. Zheng, S. Qiu, C. Shi, Q. Ma, 2024 https://scholar.google.com/scholar?q=Towards+Lifelong+Learning+of+Large+Language+Models%3A+A+Survey 7. Loss of Plasticity in Deep Continual Learning — S. Dohare, J. F. Hernandez-Garcia, Q. Lan, P. Rahman, A. R. Mahmood, R. S. Sutton, 2024 https://scholar.google.com/scholar?q=Loss+of+Plasticity+in+Deep+Continual+Learning 8. A Survey on Large Language Model Based Autonomous Agents — L. Wang, C. Ma, X. Feng, et al., 2024 https://scholar.google.com/scholar?q=A+Survey+on+Large+Language+Model+Based+Autonomous+Agents 9. WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language Models — P. Wang, Z. Li, N. Zhang, Z. Xu, Y. Yao, Y. Jiang, P. Xie, F. Huang, H. Chen, 2024 https://scholar.google.com/scholar?q=WISE%3A+Rethinking+the+Knowledge+Memory+for+Lifelong+Model+Editing+of+Large+Language+Models 10. Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem — M. McCloskey, N. J. Cohen, 1989 https://scholar.google.com/scholar?q=Catastrophic+Interference+in+Connectionist+Networks%3A+The+Sequential+Learning+Problem Interactive Visualization: Adapting Without Forgetting: A Lifelong Learning Roadmap for LLM Agents
-
664
Distributed Weight Data Parallelism Cuts LLM Inference Stalls
This episode explores DWDP (Distributed Weight Data Parallelism), a new NVIDIA-authored approach to LLM inference on NVL72 systems that targets a subtle but costly inefficiency: GPUs sitting idle while they wait to synchronize with slower peers. The hosts unpack how existing model-parallelism strategies—expert, tensor, and pipeline parallelism—all share a hidden flaw, forcing every GPU to hit a synchronization barrier at each layer boundary, which the paper's own baseline shows can waste around twelve percent of total inference time even under ordinary workload imbalance. They explain why smarter scheduling alone (cache-aware or load-aware routing) can't fix this, since it only shrinks the imbalance feeding into the wait rather than eliminating the wait itself. The discussion then turns to DWDP's core idea: keeping GPUs fully data-parallel while having each one asynchronously prefetch missing expert weights from peers on demand, timed to hide the fetch behind ongoing compute. Listeners interested in the mechanics of large-scale MoE inference, GPU synchronization bottlenecks, and practical systems-level solutions to straggler problems will find the technical walkthrough especially rewarding. Sources: 1. DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72 — Wanqian Li, Jintao Peng, Zongfei Jing, Tianyu Zhang, Ze Long, Xianjie Qiao, Xiaoming Chen, Dongxu Yang, Kefeng Duan, June Yang, 2026 http://arxiv.org/abs/2604.01621 2. DeepSeek-V3 Technical Report — DeepSeek-AI, Aixin Liu, Bei Feng, et al., 2024 https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report 3. Mooncake: Trading More Storage for Less Computation — A KVCache-centric Architecture for Serving LLM Chatbot — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, et al., 2025 https://scholar.google.com/scholar?q=Mooncake%3A+Trading+More+Storage+for+Less+Computation+%E2%80%94+A+KVCache-centric+Architecture+for+Serving+LLM+Chatbot 4. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, et al., 2024 https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving 5. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, et al., 2024 https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting 6. Tutel: Adaptive Mixture-of-Experts at Scale — Changho Hwang, Wei Cui, Yifan Xiong, et al., 2023 https://scholar.google.com/scholar?q=Tutel%3A+Adaptive+Mixture-of-Experts+at+Scale 7. DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale — Samyam Rajbhandari, Conglong Li, Zhewei Yao, et al., 2022 https://scholar.google.com/scholar?q=DeepSpeed-MoE%3A+Advancing+Mixture-of-Experts+Inference+and+Training+to+Power+Next-Generation+AI+Scale Interactive Visualization: Distributed Weight Data Parallelism Cuts LLM Inference Stalls
-
663
Cross-Family Speculative Prefill Cuts Long-Context Latency
This episode explores cross-family speculative prefill, a technique for cutting long-context inference latency by using a small "draft" model to identify which parts of a lengthy prompt matter before a much larger target model processes it. The hosts unpack why this is a hard problem in principle — draft and target models often use completely different tokenizers and architectures, meaning attention-based importance signals shouldn't obviously transfer between them — and trace the lineage from speculative decoding through the original same-family Speculative Prefill work to this paper's cross-family generalization. They highlight the practical motivation: models like DeepSeek and Kimi-K2 have no smaller sibling in their own family, so a technique that only works with matched draft/target pairs is a dead end for real deployments. Key results discussed include an 18x reduction in time-to-first-token, and the episode weighs supporting evidence from prior work on attention sinks against the stronger, less obvious claim that a full salience ranking over a 100,000-token document can transfer across unrelated architectures. Listeners interested in practical LLM efficiency techniques and the mechanics of long-context inference will find the back-and-forth skepticism over whether the method should even work, given the tokenizer mismatch, particularly engaging. Sources: 1. Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Models — Shubhangi Upasani, Ravi Shanker Raju, Bo Li, Mengmeng Ji, John Long, Chen Wu, Urmish Thakker, Guangtao Wang, 2026 http://arxiv.org/abs/2603.02631 2. Speculative Prefill — Liu et al., 2025 https://scholar.google.com/scholar?q=Speculative+Prefill 3. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023 https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding 4. Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023 (ICLR 2024) https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks 5. SnapKV: LLM Knows What You Are Looking For Before Generation — Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, Deming Chen, 2024 https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+Are+Looking+For+Before+Generation 6. LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models — Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, Lili Qiu, 2023 (EMNLP 2023) https://scholar.google.com/scholar?q=LLMLingua%3A+Compressing+Prompts+for+Accelerated+Inference+of+Large+Language+Models 7. LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression — Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, Lili Qiu, 2024 (ACL 2024) https://scholar.google.com/scholar?q=LongLLMLingua%3A+Accelerating+and+Enhancing+LLMs+in+Long+Context+Scenarios+via+Prompt+Compression 8. LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression — Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Ruhle, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, Dongmei Zhang, 2024 (ACL Findings 2024) https://scholar.google.com/scholar?q=LLMLingua-2%3A+Data+Distillation+for+Efficient+and+Faithful+Task-Agnostic+Prompt+Compression 9. Learning to Compress Prompts with Gist Tokens — Jesse Mu, Xiang Lisa Li, Noah Goodman, 2023 (NeurIPS 2023) https://scholar.google.com/scholar?q=Learning+to+Compress+Prompts+with+Gist+Tokens 10. Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation — Jingyu Liu, Beidi Chen, Ce Zhang, 2025 https://scholar.google.com/scholar?q=Speculative+Prefill%3A+Turbocharging+TTFT+with+Lightweight+and+Training-Free+Token+Importance+Estimation 11. SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators — Jonathan Li, Nasim Farahini, et al. (SambaNova), 2025 https://scholar.google.com/scholar?q=SnapStream%3A+Efficient+Long+Sequence+Decoding+on+Dataflow+Accelerators 12. LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference — Guangtao Wang, Shubhangi Upasani, Chen Wu, et al. (SambaNova), 2025 https://scholar.google.com/scholar?q=LLMs+Know+What+to+Drop%3A+Self-Attention+Guided+KV+Cache+Eviction+for+Efficient+Long-Context+Inference 13. MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention — Huiqiang Jiang, Yucheng Li, Chengruidong Zhang, et al., 2024 https://scholar.google.com/scholar?q=MInference+1.0%3A+Accelerating+Pre-filling+for+Long-Context+LLMs+via+Dynamic+Sparse+Attention Interactive Visualization: Cross-Family Speculative Prefill Cuts Long-Context Latency
-
662
AdaJEPA: Self-Adapting Latent World Models via Test-Time MPC
This episode explores AdaJEPA, an adaptive latent world model that challenges the standard "train once, freeze forever" assumption behind robot planning systems. The hosts trace the technical lineage from Yann LeCun's Joint-Embedding Predictive Architecture concept through model predictive control's decades-old roots in process engineering and rocket landing, showing how these pieces combine to let a deployed robot keep updating its internal model using only the consequences of its own actions — no new labels, demonstrations, or retraining pipeline required. Central to the discussion is how distribution shift causes small prediction errors to compound across multi-step planning horizons, and how test-time adaptation, borrowed from image classification and paralleled to cerebellar motor learning, closes that loop by treating each observed transition as a live training example. The conversation grounds abstract control theory in concrete deployment scenarios, from unfamiliar object shapes to shifting friction and lighting. Listeners interested in robotics, control theory, or self-supervised learning will find a clear walkthrough of why frozen world models fail in the wild and what it means for a model to keep learning after "training" officially ends. Sources: 1. AdaJEPA: An Adaptive Latent World Model — Ying Wang, Oumayma Bounou, Yann LeCun, Mengye Ren, 2026 http://arxiv.org/abs/2606.32026 2. Model Predictive Control: Theory and Practice — A Survey — Carlos E. García, David M. Prett, Manfred Morari, 1989 https://scholar.google.com/scholar?q=Model+Predictive+Control%3A+Theory+and+Practice+%E2%80%94+A+Survey 3. Model Predictive Control: Classical, Robust and Stochastic — Basil Kouvaritakis, Mark Cannon, 2016 https://scholar.google.com/scholar?q=Model+Predictive+Control%3A+Classical%2C+Robust+and+Stochastic 4. Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models (PETS) — Kurtland Chua, Roberto Calandra, Rowan McAllister, Sergey Levine, 2018 https://scholar.google.com/scholar?q=Deep+Reinforcement+Learning+in+a+Handful+of+Trials+using+Probabilistic+Dynamics+Models+%28PETS%29 5. Learning Latent Dynamics for Planning from Pixels (PlaNet) — Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, James Davidson, 2019 https://scholar.google.com/scholar?q=Learning+Latent+Dynamics+for+Planning+from+Pixels+%28PlaNet%29 6. Dino-wm: World models on pre-trained visual features enable zero-shot planning — Zhou, G., Pan, H., LeCun, Y., and Pinto, L., 2025 https://scholar.google.com/scholar?q=Dino-wm%3A+World+models+on+pre-trained+visual+features+enable+zero-shot+planning 7. Temporal straightening for latent planning — Wang, Y., Bounou, O., Zhou, G., Balestriero, R., Rudner, T. G., LeCun, Y., and Ren, M., 2026 https://scholar.google.com/scholar?q=Temporal+straightening+for+latent+planning 8. Closing the train-test gap in world models for gradient-based planning — Parthasarathy, A., Kalra, N., Agrawal, R., LeCun, Y., Bounou, O., Izmailov, P., and Goldblum, M., 2025 https://scholar.google.com/scholar?q=Closing+the+train-test+gap+in+world+models+for+gradient-based+planning 9. Td-mpc2: Scalable, robust world models for continuous control — Hansen, N., Su, H., and Wang, X., 2024 https://scholar.google.com/scholar?q=Td-mpc2%3A+Scalable%2C+robust+world+models+for+continuous+control 10. Adawm: Adaptive world model based planning for autonomous driving — Wang, H., Ye, X., Tao, F., Pan, C., Mallik, A., Yaman, B., Ren, L., and Zhang, J., 2025 https://scholar.google.com/scholar?q=Adawm%3A+Adaptive+world+model+based+planning+for+autonomous+driving 11. Test-time training with self-supervision for generalization under distribution shifts — Sun, Y., Wang, X., Liu, Z., Miller, J., Efros, A. A., and Hardt, M., 2020 https://scholar.google.com/scholar?q=Test-time+training+with+self-supervision+for+generalization+under+distribution+shifts Interactive Visualization: AdaJEPA: Self-Adapting Latent World Models via Test-Time MPC
-
661
Test-Time Training Turns EDA Feedback Into Live Weight Updates for RTL
This episode explores Alpha-RTL, a framework applying test-time training to RTL hardware optimization, where an LLM updates its own weights live for each chip design using real EDA toolchain feedback rather than a static, pre-trained policy. The discussion contrasts this approach with two existing camps: agentic search methods (like REvolution) that iterate over a frozen model and discard synthesis feedback after each run, and training-time reinforcement learning (like ChipSeek) that learns once offline and only samples at inference. It unpacks why functional correctness in Verilog is a weak proxy for what chip teams actually optimize — PPA, the area-delay-power product measured only after synthesis — and traces the paper's core techniques back to their origins: test-time training from Sun et al.'s 2020 UC Berkeley work, and PUCT search from Kocsis and Szepesvári's 2006 UCT paper, extended here into a persistent state pool of Verilog candidates refined over gradient updates rather than resampled from scratch. Listeners interested in the mechanics of closing the loop between LLM code generation and physical design constraints — and the unusual tradeoff of burning GPU-hours to fine-tune a model for a single, disposable hardware block — will find the episode's breakdown of RLVR-style staged verification (compile, simulate, synthesize) particularly useful. Sources: 1. Alpha-RTL: Test-Time Training for RTL Hardware Optimization — Peilong Zhou, Zhirong Chen, Cangyuan Li, Haoyu Gao, Kaiyan Chang, Ziming Qu, Ying Wang, 2026 http://arxiv.org/abs/2606.05253 2. Bandit based Monte-Carlo Planning — Levente Kocsis, Csaba Szepesvári, 2006 https://scholar.google.com/scholar?q=Bandit+based+Monte-Carlo+Planning 3. A Survey of Monte Carlo Tree Search Methods — Cameron Browne, Edward Powley, Daniel Whitehouse, Simon Lucas, Peter Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, Simon Colton, 2012 https://scholar.google.com/scholar?q=A+Survey+of+Monte+Carlo+Tree+Search+Methods 4. Mastering the game of Go with deep neural networks and tree search — David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, et al. (DeepMind), 2016 https://scholar.google.com/scholar?q=Mastering+the+game+of+Go+with+deep+neural+networks+and+tree+search 5. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play — David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, et al. (DeepMind), 2018 https://scholar.google.com/scholar?q=A+general+reinforcement+learning+algorithm+that+masters+chess%2C+shogi%2C+and+Go+through+self-play 6. Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm (AlphaZero) — David Silver et al., 2017/2018 https://scholar.google.com/scholar?q=Mastering+Chess+and+Shogi+by+Self-Play+with+a+General+Reinforcement+Learning+Algorithm+%28AlphaZero%29 7. A Graph Placement Methodology for Fast Chip Design (AlphaChip) — Azalia Mirhoseini, Anna Goldie, Mustafa Yazgan, et al., 2021 https://scholar.google.com/scholar?q=A+Graph+Placement+Methodology+for+Fast+Chip+Design+%28AlphaChip%29 8. Data-Driven Offline Optimization for Architecting Hardware Accelerators (PRIME) — Aviral Kumar, Amir Yazdanbakhsh, Milad Hashemi, Kevin Swersky, Sergey Levine, 2021/2022 https://scholar.google.com/scholar?q=Data-Driven+Offline+Optimization+for+Architecting+Hardware+Accelerators+%28PRIME%29 9. SymbiYosys / eqy (formal equivalence checking for Yosys-based flows) — YosysHQ / Claire Wolf and contributors, ongoing https://scholar.google.com/scholar?q=SymbiYosys+%2F+eqy+%28formal+equivalence+checking+for+Yosys-based+flows%29
-
660
Main Trust Issue in FPGA HLS Design Workflow
This episode examines ContractHIL-HLS, a paper from Jingbo Zhang and colleagues at Beijing University of Technology (posted to arXiv July 28, 2026) that tackles high-level synthesis for FPGA design, where LLM-generated hardware can compile cleanly and pass simulation yet still fail on real silicon due to timing violations, routing congestion, or power overruns that only surface during actual synthesis and place-and-route. The discussion contrasts this work with prior efforts like Chip-Chat, RTLLM, and HLS-Eval, arguing those prove models can generate hardware code but not that a workflow can reliably preserve design intent and incorporate tool feedback across multiple steps. Rather than relying on conversational role-prompting, where constraints can silently drift or vanish between turns, the paper's architecture splits agents by transformation type — a Contract Agent converts natural language into a structured object with named fields for interface, constraints, and validation policy, an HTML Agent renders it stably, and a Hardware-in-the-Loop Agent implements and revises designs using real Vitis HLS synthesis, Vivado place-and-route, and board bring-up rather than trusting the model's own claims. The hosts debate whether structured fields actually prevent drift better than conversational memory does, landing on the distinction that a missing field is inspectable while conversational drift is not, though enforcement remains an open question. Listeners interested in how hardware-design automation might borrow validation rigor from aerospace and control-systems engineering will find the explanation of Hardware-in-the-Loop testing, and its adaptation to catch AI-generated designs before they reach costly physical fabrication, especially compelling. Sources: 1. ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design — Jingbo Zhang, Haoxiang Sun, Wenbo Wang, Wenbo Zhang, 2026 http://arxiv.org/abs/2607.25283 2. LegUp: High-Level Synthesis for FPGA-Based Processor/Accelerator Systems — Andrew Canis, Jongsok Choi, Mark Aldham, Victor Zhang, Ahmed Kammoona, Jason Anderson, Stephen Brown, Tomasz Czajkowski, 2011 (FPGA conference; extended in ACM TODAES 2013) https://scholar.google.com/scholar?q=LegUp%3A+High-Level+Synthesis+for+FPGA-Based+Processor%2FAccelerator+Systems 3. Fast Inference of Deep Neural Networks in FPGAs for Particle Physics (hls4ml) — Javier Duarte, Song Han, Philip Harris, et al., 2018 https://scholar.google.com/scholar?q=Fast+Inference+of+Deep+Neural+Networks+in+FPGAs+for+Particle+Physics+%28hls4ml%29 4. Chip-Chat: Challenges and Opportunities in Conversational Hardware Design — Jason Blocklove, Siddharth Garg, Ramesh Karri, Hammond Pearce, 2023 https://scholar.google.com/scholar?q=Chip-Chat%3A+Challenges+and+Opportunities+in+Conversational+Hardware+Design 5. AutoChip: Automating HDL Generation Using LLM Feedback — Shailja Thakur et al., 2023 https://scholar.google.com/scholar?q=AutoChip%3A+Automating+HDL+Generation+Using+LLM+Feedback 6. MnasNet: Platform-Aware Neural Architecture Search for Mobile — Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, Quoc V. Le, 2019 https://scholar.google.com/scholar?q=MnasNet%3A+Platform-Aware+Neural+Architecture+Search+for+Mobile 7. FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search — Bichen Wu, Xiaoliang Zhang, Kaiwen Weng, Yandong Guo, Peizhao Zhang, Yanghan Wang, Kurt Keutzer, Peter Vajda, 2019 https://scholar.google.com/scholar?q=FBNet%3A+Hardware-Aware+Efficient+ConvNet+Design+via+Differentiable+Neural+Architecture+Search 8. Learning Dexterous In-Hand Manipulation — OpenAI (Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, et al.), 2020 (IJRR; earlier preprint 2018) https://scholar.google.com/scholar?q=Learning+Dexterous+In-Hand+Manipulation 9. CRYSTALS-Kyber: A CCA-Secure Module-Lattice-Based KEM — Joppe Bos, Léo Ducas, Eike Kiltz, Tancrède Lepoint, Vadim Lyubashevsky, John M. Schanck, Peter Schwabe, Gregor Seiler, Damien Stehlé, 2018 https://scholar.google.com/scholar?q=CRYSTALS-Kyber%3A+A+CCA-Secure+Module-Lattice-Based+KEM 10. A Compact Hardware Implementation of CCA-Secure Key Exchange Mechanism CRYSTALS-KYBER on FPGA — Yufei Xing, Shuguo Li, 2021 https://scholar.google.com/scholar?q=A+Compact+Hardware+Implementation+of+CCA-Secure+Key+Exchange+Mechanism+CRYSTALS-KYBER+on+FPGA 11. Module-Lattice-Based Key-Encapsulation Mechanism Standard (FIPS 203) — National Institute of Standards and Technology (NIST), 2024 https://scholar.google.com/scholar?q=Module-Lattice-Based+Key-Encapsulation+Mechanism+Standard+%28FIPS+203%29 12. KyberMat and CRYPHTOR (accelerator designs cited directly in the ContractHIL-HLS paper) — Not independently verified here — cited by the ContractHIL-HLS authors as references [8] and [9], Recent (post-2023, exact years unconfirmed) https://scholar.google.com/scholar?q=KyberMat+and+CRYPHTOR+%28accelerator+designs+cited+directly+in+the+ContractHIL-HLS+paper%29 13. HLS-Eval: A Benchmark and Framework for Evaluating LLMs on High-Level Synthesis Design Tasks — S. Abi-Karam, C. Hao, 2025 https://scholar.google.com/scholar?q=HLS-Eval%3A+A+Benchmark+and+Framework+for+Evaluating+LLMs+on+High-Level+Synthesis+Design+Tasks 14. SAGE-HLS: Syntax-Aware AST-Guided LLM for High-Level Synthesis Code Generation — M. Z. S. Khan, N. Mashnoor, M. Akyash et al., 2025 https://scholar.google.com/scholar?q=SAGE-HLS%3A+Syntax-Aware+AST-Guided+LLM+for+High-Level+Synthesis+Code+Generation 15. A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT — J. White, Q. Fu, S. Hays et al., 2023 https://scholar.google.com/scholar?q=A+Prompt+Pattern+Catalog+to+Enhance+Prompt+Engineering+with+ChatGPT 16. Evaluating Large Language Models Trained on Code — M. Chen, J. Tworek, H. Jun et al. (OpenAI Codex/HumanEval), 2021 https://scholar.google.com/scholar?q=Evaluating+Large+Language+Models+Trained+on+Code 17. KyberMat: Efficient Accelerator for Matrix-Vector Polynomial Multiplication in CRYSTALS-Kyber via NTT and Polyphase Decomposition — W. Tan, Y. Lao, K. K. Parhi, 2023 https://scholar.google.com/scholar?q=KyberMat%3A+Efficient+Accelerator+for+Matrix-Vector+Polynomial+Multiplication+in+CRYSTALS-Kyber+via+NTT+and+Polyphase+Decomposition Interactive Visualization: Main Trust Issue in FPGA HLS Design Workflow
-
659
The Unlearnability Phenomenon in RLVR Reasoning Models
This episode explores "The Unlearnability Phenomenon in RLVR for Language Models" by Yulin Chen and colleagues at NYU, which uncovers a puzzling failure mode in reinforcement learning with verifiable reward (RLVR)—the training method underlying reasoning models like o1, o3, DeepSeek-R1, and QwQ. The hosts unpack how GRPO, the algorithm popularized by DeepSeek, relies on reward variance across sampled rollouts to compute learning signals, and how the paper's authors tracked individual hard training examples to discover that some receive genuine positive reward repeatedly yet never show improved success rates—even after training converges. The discussion probes why this defies basic policy-gradient intuition, since a rewarded rollout should become more probable regardless of whether the model got the right answer through skill or luck. The core investigative thread centers on gradient cosine similarity—checking whether an example's own learning signal aligns with or fights against the rest of the training batch—as the lens for explaining why some correctly-solved problems never stick. Listeners interested in the mechanics and hidden limits of frontier reasoning-model training will find this a sharp look at a ceiling effect invisible in ordinary loss curves. Sources: 1. The Unlearnability Phenomenon in RLVR for Language Models — Yulin Chen, He He, Chen Zhao, 2026 http://arxiv.org/abs/2605.16787 2. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models — Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y.K. Li, Y. Wu, Daya Guo (DeepSeek-AI), 2024 https://scholar.google.com/scholar?q=DeepSeekMath%3A+Pushing+the+Limits+of+Mathematical+Reasoning+in+Open+Language+Models 3. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI (Daya Guo et al.), 2025 https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning 4. DAPO: An Open-Source LLM Reinforcement Learning System at Scale — Qiying Yu, Zheng Zhang, Yu Yue, Mingxuan Wang, et al. (ByteDance Seed / Tsinghua AIR), 2025 https://scholar.google.com/scholar?q=DAPO%3A+An+Open-Source+LLM+Reinforcement+Learning+System+at+Scale 5. Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? — Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, et al. (Tsinghua University), 2025 https://scholar.google.com/scholar?q=Does+Reinforcement+Learning+Really+Incentivize+Reasoning+Capacity+in+LLMs+Beyond+the+Base+Model%3F 6. OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling — Z. Wang, F. Zhou, X. Li, P. Liu, 2025 https://scholar.google.com/scholar?q=OctoThinker%3A+Mid-training+Incentivizes+Reinforcement+Learning+Scaling 7. Arithmetic without Algorithms: Language Models Solve Math with a Bag of Heuristics — Y. Nikankin, A. Reusch, A. Mueller, Y. Belinkov, 2025 (ICLR) https://scholar.google.com/scholar?q=Arithmetic+without+Algorithms%3A+Language+Models+Solve+Math+with+a+Bag+of+Heuristics 8. Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification — A. Zhang, Y. Chen, J. Pan, C. Zhao, A. Panda, J. Li, H. He, 2025 https://scholar.google.com/scholar?q=Reasoning+Models+Know+When+They%27re+Right%3A+Probing+Hidden+States+for+Self-Verification 9. The Invisible Leash: Why RLVR May or May Not Escape Its Origin — F. Wu, W. Xuan, X. Lu, M. Liu, Y. Dong, Z. Harchaoui, Y. Choi, 2026 https://scholar.google.com/scholar?q=The+Invisible+Leash%3A+Why+RLVR+May+or+May+Not+Escape+Its+Origin 10. The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models — G. Cui, Y. Zhang, J. Chen, L. Yuan, Z. Wang, Y. Zuo, et al., 2025 https://scholar.google.com/scholar?q=The+Entropy+Mechanism+of+Reinforcement+Learning+for+Reasoning+Language+Models Interactive Visualization: The Unlearnability Phenomenon in RLVR Reasoning Models
We're indexing this podcast's transcripts for the first time — this can take a minute or two. We'll show results as soon as they're ready.
No matches for "" in this podcast's transcripts.
No topics indexed yet for this podcast.
Loading reviews...
ABOUT THIS SHOW
AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.
HOSTED BY
mcgrof
CATEGORIES
Loading similar podcasts...