EPISODE · Jul 31, 2026
The Unlearnability Phenomenon in RLVR Reasoning Models
from AI Post Transformers
This episode explores "The Unlearnability Phenomenon in RLVR for Language Models" by Yulin Chen and colleagues at NYU, which uncovers a puzzling failure mode in reinforcement learning with verifiable reward (RLVR)—the training method underlying reasoning models like o1, o3, DeepSeek-R1, and QwQ. The hosts unpack how GRPO, the algorithm popularized by DeepSeek, relies on reward variance across sampled rollouts to compute learning signals, and how the paper's authors tracked individual hard training examples to discover that some receive genuine positive reward repeatedly yet never show improved success rates—even after training converges. The discussion probes why this defies basic policy-gradient intuition, since a rewarded rollout should become more probable regardless of whether the model got the right answer through skill or luck. The core investigative thread centers on gradient cosine similarity—checking whether an example's own learning signal aligns with or fights against the rest of the training batch—as the lens for explaining why some correctly-solved problems never stick. Listeners interested in the mechanics and hidden limits of frontier reasoning-model training will find this a sharp look at a ceiling effect invisible in ordinary loss curves. Sources: 1. The Unlearnability Phenomenon in RLVR for Language Models — Yulin Chen, He He, Chen Zhao, 2026 http://arxiv.org/abs/2605.16787 2. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models — Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y.K. Li, Y. Wu, Daya Guo (DeepSeek-AI), 2024 https://scholar.google.com/scholar?q=DeepSeekMath%3A+Pushing+the+Limits+of+Mathematical+Reasoning+in+Open+Language+Models 3. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI (Daya Guo et al.), 2025 https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning 4. DAPO: An Open-Source LLM Reinforcement Learning System at Scale — Qiying Yu, Zheng Zhang, Yu Yue, Mingxuan Wang, et al. (ByteDance Seed / Tsinghua AIR), 2025 https://scholar.google.com/scholar?q=DAPO%3A+An+Open-Source+LLM+Reinforcement+Learning+System+at+Scale 5. Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? — Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, et al. (Tsinghua University), 2025 https://scholar.google.com/scholar?q=Does+Reinforcement+Learning+Really+Incentivize+Reasoning+Capacity+in+LLMs+Beyond+the+Base+Model%3F 6. OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling — Z. Wang, F. Zhou, X. Li, P. Liu, 2025 https://scholar.google.com/scholar?q=OctoThinker%3A+Mid-training+Incentivizes+Reinforcement+Learning+Scaling 7. Arithmetic without Algorithms: Language Models Solve Math with a Bag of Heuristics — Y. Nikankin, A. Reusch, A. Mueller, Y. Belinkov, 2025 (ICLR) https://scholar.google.com/scholar?q=Arithmetic+without+Algorithms%3A+Language+Models+Solve+Math+with+a+Bag+of+Heuristics 8. Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification — A. Zhang, Y. Chen, J. Pan, C. Zhao, A. Panda, J. Li, H. He, 2025 https://scholar.google.com/scholar?q=Reasoning+Models+Know+When+They%27re+Right%3A+Probing+Hidden+States+for+Self-Verification 9. The Invisible Leash: Why RLVR May or May Not Escape Its Origin — F. Wu, W. Xuan, X. Lu, M. Liu, Y. Dong, Z. Harchaoui, Y. Choi, 2026 https://scholar.google.com/scholar?q=The+Invisible+Leash%3A+Why+RLVR+May+or+May+Not+Escape+Its+Origin 10. The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models — G. Cui, Y. Zhang, J. Chen, L. Yuan, Z. Wang, Y. Zuo, et al., 2025 https://scholar.google.com/scholar?q=The+Entropy+Mechanism+of+Reinforcement+Learning+for+Reasoning+Language+Models Interactive Visualization: The Unlearnability Phenomenon in RLVR Reasoning Models
Embed this episode
NOW PLAYING
The Unlearnability Phenomenon in RLVR Reasoning Models
No transcript for this episode yet
Similar Episodes
No similar episodes found.