Weak-to-Strong On-Policy Distillation Beats the Teacher episode artwork

EPISODE · Aug 28, 2026

Weak-to-Strong On-Policy Distillation Beats the Teacher

from AI Post Transformers

This episode examines Weak-to-Strong On-Policy Distillation, a paper from University of Maryland, Microsoft Research, and MBZUAI showing that an 8-billion-parameter student model can outperform every teacher used to train it on math benchmarks. The discussion traces the method's lineage from Hinton's original 2015 knowledge distillation through DAgger's 2011 on-policy correction idea to 2023 weak-to-strong generalization work from OpenAI's Superalignment team, explaining how the paper inverts DAgger's core assumption by using a supervisor weaker than the student rather than a trusted expert. It also contrasts this approach with reinforcement learning from verifiable rewards, which gives only a sparse end-of-rollout signal, versus on-policy distillation's dense per-token feedback on the student's own generated trajectories. The hosts situate the work against industry precedents like Qwen3's large-to-small distillation and multi-teacher on-policy distillation, both of which still require a teacher at least as capable as the student — a constraint this paper's method aims to eliminate. Listeners interested in how weaker models might keep training stronger ones as the field runs out of better supervisors will find the mechanism and its research lineage laid out in detail. Sources: 1. Weak-to-Strong On-Policy Distillation — Fangxu Yu, Weijia Xu, Michael Xu, Tianyi Zhou, Zinan Lin, 2026 http://arxiv.org/abs/2607.26246 2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015 https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network 3. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29 4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (Generalized Knowledge Distillation, GKD) — Rishabh Agarwal, Nino Vieillard, et al. (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes+%28Generalized+Knowledge+Distillation%2C+GKD%29 5. On-Policy Distillation (blog post / technical report) — Thinking Machines Lab (cited in the paper as 'Lu & Lab'), 2025 https://scholar.google.com/scholar?q=On-Policy+Distillation+%28blog+post+%2F+technical+report%29 6. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision — Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, Jeff Wu, 2023 https://scholar.google.com/scholar?q=Weak-to-Strong+Generalization%3A+Eliciting+Strong+Capabilities+With+Weak+Supervision 7. Weak-to-Strong Generalization beyond Accuracy: a Roadmap in Codegen, Safety, and Beyond (or closely related 2024 weak-to-strong follow-up work, cited in the paper as 'Ning et al., 2024') — Ning et al., 2024 https://scholar.google.com/scholar?q=Weak-to-Strong+Generalization+beyond+Accuracy%3A+a+Roadmap+in+Codegen%2C+Safety%2C+and+Beyond+%28or+closely+related+2024+weak-to-strong+follow-up+work%2C+cited+in+the+paper+as+%27Ning+et+al.%2C+2024%27%29 8. Illustrating Reinforcement Learning from Human Feedback / scalable oversight lineage (e.g. Christiano et al., 'Deep Reinforcement Learning from Human Preferences', and related scalable-oversight work) — Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, Dario Amodei (RLHF); related scalable oversight literature, 2017 (RLHF) / ongoing scalable oversight literature https://scholar.google.com/scholar?q=Illustrating+Reinforcement+Learning+from+Human+Feedback+%2F+scalable+oversight+lineage+%28e.g.+Christiano+et+al.%2C+%27Deep+Reinforcement+Learning+from+Human+Preferences%27%2C+and+related+scalable-oversight+work%29 9. Contrastive Decoding: Open-ended Text Generation as Optimization — Li, Holtzman, Fried, Liang, Eisner, Hashimoto, Zettlemoyer, Lewis, 2023 https://scholar.google.com/scholar?q=Contrastive+Decoding%3A+Open-ended+Text+Generation+as+Optimization 10. Distillation Scaling Laws — Busbridge, Ramapuram, Ablin, et al. (Apple), 2024 https://scholar.google.com/scholar?q=Distillation+Scaling+Laws 11. Speculative Decoding papers (e.g., Leviathan et al., 'Fast Inference from Transformers via Speculative Decoding', 2023) — Leviathan, Kalman, Matias, 2023 https://scholar.google.com/scholar?q=Speculative+Decoding+papers+%28e.g.%2C+Leviathan+et+al.%2C+%27Fast+Inference+from+Transformers+via+Speculative+Decoding%27%2C+2023%29 Interactive Visualization: Weak-to-Strong On-Policy Distillation Beats the Teacher

Episode metadata supplied by the publisher feed · Published Aug 28, 2026

Embed this episode

NOW PLAYING

Weak-to-Strong On-Policy Distillation Beats the Teacher

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on August 28, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!