EPISODE · Aug 28, 2026
Weak-to-Strong On-Policy Distillation Beats the Teacher
from AI Post Transformers
This episode examines Weak-to-Strong On-Policy Distillation, a paper from University of Maryland, Microsoft Research, and MBZUAI showing that an 8-billion-parameter student model can outperform every teacher used to train it on math benchmarks. The discussion traces the method's lineage from Hinton's original 2015 knowledge distillation through DAgger's 2011 on-policy correction idea to 2023 weak-to-strong generalization work from OpenAI's Superalignment team, explaining how the paper inverts DAgger's core assumption by using a supervisor weaker than the student rather than a trusted expert. It also contrasts this approach with reinforcement learning from verifiable rewards, which gives only a sparse end-of-rollout signal, versus on-policy distillation's dense per-token feedback on the student's own generated trajectories. The hosts situate the work against industry precedents like Qwen3's large-to-small distillation and multi-teacher on-policy distillation, both of which still require a teacher at least as capable as the student — a constraint this paper's method aims to eliminate. Listeners interested in how weaker models might keep training stronger ones as the field runs out of better supervisors will find the mechanism and its research lineage laid out in detail. Sources: 1. Weak-to-Strong On-Policy Distillation — Fangxu Yu, Weijia Xu, Michael Xu, Tianyi Zhou, Zinan Lin, 2026 http://arxiv.org/abs/2607.26246 2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015 https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network 3. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29 4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (Generalized Knowledge Distillation, GKD) — Rishabh Agarwal, Nino Vieillard, et al. (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes+%28Generalized+Knowledge+Distillation%2C+GKD%29 5. On-Policy Distillation (blog post / technical report) — Thinking Machines Lab (cited in the paper as 'Lu & Lab'), 2025 https://scholar.google.com/scholar?q=On-Policy+Distillation+%28blog+post+%2F+technical+report%29 6. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision — Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, Jeff Wu, 2023 https://scholar.google.com/scholar?q=Weak-to-Strong+Generalization%3A+Eliciting+Strong+Capabilities+With+Weak+Supervision 7. Weak-to-Strong Generalization beyond Accuracy: a Roadmap in Codegen, Safety, and Beyond (or closely related 2024 weak-to-strong follow-up work, cited in the paper as 'Ning et al., 2024') — Ning et al., 2024 https://scholar.google.com/scholar?q=Weak-to-Strong+Generalization+beyond+Accuracy%3A+a+Roadmap+in+Codegen%2C+Safety%2C+and+Beyond+%28or+closely+related+2024+weak-to-strong+follow-up+work%2C+cited+in+the+paper+as+%27Ning+et+al.%2C+2024%27%29 8. Illustrating Reinforcement Learning from Human Feedback / scalable oversight lineage (e.g. Christiano et al., 'Deep Reinforcement Learning from Human Preferences', and related scalable-oversight work) — Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, Dario Amodei (RLHF); related scalable oversight literature, 2017 (RLHF) / ongoing scalable oversight literature https://scholar.google.com/scholar?q=Illustrating+Reinforcement+Learning+from+Human+Feedback+%2F+scalable+oversight+lineage+%28e.g.+Christiano+et+al.%2C+%27Deep+Reinforcement+Learning+from+Human+Preferences%27%2C+and+related+scalable-oversight+work%29 9. Contrastive Decoding: Open-ended Text Generation as Optimization — Li, Holtzman, Fried, Liang, Eisner, Hashimoto, Zettlemoyer, Lewis, 2023 https://scholar.google.com/scholar?q=Contrastive+Decoding%3A+Open-ended+Text+Generation+as+Optimization 10. Distillation Scaling Laws — Busbridge, Ramapuram, Ablin, et al. (Apple), 2024 https://scholar.google.com/scholar?q=Distillation+Scaling+Laws 11. Speculative Decoding papers (e.g., Leviathan et al., 'Fast Inference from Transformers via Speculative Decoding', 2023) — Leviathan, Kalman, Matias, 2023 https://scholar.google.com/scholar?q=Speculative+Decoding+papers+%28e.g.%2C+Leviathan+et+al.%2C+%27Fast+Inference+from+Transformers+via+Speculative+Decoding%27%2C+2023%29 Interactive Visualization: Weak-to-Strong On-Policy Distillation Beats the Teacher
Embed this episode
NOW PLAYING
Weak-to-Strong On-Policy Distillation Beats the Teacher
No transcript for this episode yet
Similar Episodes
No similar episodes found.