EPISODE · Aug 15, 2026
Learning, Fast and Slow: LLMs That Adapt Without Forgetting
from AI Post Transformers
This episode explores catastrophic forgetting and plasticity loss in RL-trained language models, and introduces "Fast-Slow Training," a method combining slow weight updates (RLVR) with fast in-context learning to address both. The hosts unpack the distinction between RLVR's automatic, verifiable rewards and traditional RLHF, then dig into two separate failure modes of pure RL post-training: models forgetting general competence while chasing a narrow reward signal, and a subtler loss of plasticity where updates leave models increasingly unable to absorb new tasks. Framing the two training channels as a System 1/System 2 split, the discussion centers on the paper's headline result — combining both channels reaches RL's peak accuracy with up to three times fewer samples, drifts up to seventy percent less from the base model, and preserves the capacity to learn subsequent tasks where pure RL stalls. Listeners interested in the mechanics and tradeoffs of continual learning in large language models will find a grounded walkthrough of why prompting alone hits a ceiling and why weight updates alone come with hidden costs. Sources: 1. Learning, Fast and Slow: Towards LLMs That Adapt Continually — Rishabh Tiwari, Kusha Sareen, Lakshya A Agrawal, Joseph E. Gonzalez, Matei Zaharia, Kurt Keutzer, Inderjit S Dhillon, Rishabh Agarwal, Devvrit Khatri, 2026 http://arxiv.org/abs/2605.12484 2. Loss of Plasticity in Deep Continual Learning — Shibhansh Dohare, J. Fernando Hernandez-Garcia, Qingfeng Lan, A. Rupam Mahmood, Richard S. Sutton, et al., 2024 (Nature; preprint circulated as 'Maintaining Plasticity via Continual Backprop' from 2021) https://scholar.google.com/scholar?q=Loss+of+Plasticity+in+Deep+Continual+Learning 3. On Warm-Starting Neural Network Training — Jordan T. Ash, Ryan P. Adams, 2020 (NeurIPS) https://scholar.google.com/scholar?q=On+Warm-Starting+Neural+Network+Training 4. The Primacy Bias in Deep Reinforcement Learning — Evgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon, Aaron Courville, 2022 (ICML) https://scholar.google.com/scholar?q=The+Primacy+Bias+in+Deep+Reinforcement+Learning 5. Understanding Plasticity in Neural Networks — Clare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Avila Pires, Razvan Pascanu, Will Dabney, 2023 (ICML) https://scholar.google.com/scholar?q=Understanding+Plasticity+in+Neural+Networks 6. RL's razor: Why online reinforcement learning forgets less — Idan Shenfeld, Jyothish Pari, Pulkit Agrawal, 2025 https://scholar.google.com/scholar?q=RL%27s+razor%3A+Why+online+reinforcement+learning+forgets+less 7. The Art of Scaling Reinforcement Learning Compute for LLMs (ScaleRL) — Devvrit Khatri, Lovish Madaan, Rishabh Tiwari, Rachit Bansal, Sai Surya Duvvuri, Manzil Zaheer, Inderjit S. Dhillon, David Brandfonbrener, Rishabh Agarwal, 2025 https://scholar.google.com/scholar?q=The+Art+of+Scaling+Reinforcement+Learning+Compute+for+LLMs+%28ScaleRL%29 8. Fine-tuning and prompt optimization: Two great steps that work better together (BetterTogether) — Dilara Soylu, Christopher Potts, Omar Khattab, 2024 https://scholar.google.com/scholar?q=Fine-tuning+and+prompt+optimization%3A+Two+great+steps+that+work+better+together+%28BetterTogether%29 9. Mitigating plasticity loss in continual reinforcement learning by reducing churn — Hongyao Tang, Johan Obando-Ceron, Pablo Samuel Castro, Aaron Courville, Glen Berseth, 2025 https://scholar.google.com/scholar?q=Mitigating+plasticity+loss+in+continual+reinforcement+learning+by+reducing+churn 10. What can you do when you have zero rewards during RL? — Jatin Prakash, Anirudh Buvanesh, 2025 https://scholar.google.com/scholar?q=What+can+you+do+when+you+have+zero+rewards+during+RL%3F Interactive Visualization: Learning, Fast and Slow: LLMs That Adapt Without Forgetting
Embed this episode
NOW PLAYING
Learning, Fast and Slow: LLMs That Adapt Without Forgetting
No transcript for this episode yet
Similar Episodes
No similar episodes found.