EPISODE · Jul 2, 2026 · 22 MIN
EP281: Restoring plasticity to over-trained AI
from Learning GenAI via SOTA Papers · host Yun Wu
Title: When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL HandoffSource: http://arxiv.org/abs/2606.09932v1Summary:This paper identifies and solves the critical 'loss of plasticity' bottleneck in the standard LLM post-training pipeline where excessive SFT inhibits subsequent RL optimization. It introduces 'Rejuvenation', a foundational training primitive that uses model fusion and neuron resets to enable robust reasoning gains during RL while preserving SFT-acquired knowledge.
Embed this episode
Ready to play
EP281: Restoring plasticity to over-trained AI
No transcript for this episode yet
Similar Episodes
No similar episodes found.