Lightning OPD: Precomputing the Teacher for Cheaper Reasoning Distillation episode artwork

EPISODE · Jul 21, 2026

Lightning OPD: Precomputing the Teacher for Cheaper Reasoning Distillation

from AI Post Transformers

This episode traces the evolution from classic knowledge distillation to on-policy distillation and finally to Lightning OPD, a technique from NVIDIA researchers for post-training large reasoning models. The hosts unpack why on-policy distillation offers denser training signal than RLVR's sparse end-of-trace rewards, but has historically required an expensive live teacher model running alongside the student throughout training. They explain the paper's key insight — that a student's rollout distribution drifts only modestly from its SFT starting point — which motivates capturing the teacher's judgments once offline rather than serving it continuously. The discussion grounds the work in its lineage, from Hinton et al.'s original 2015 distillation paper to DeepMind's 2024 Generalized Knowledge Distillation formalism, while flagging why the infrastructure savings matter even more for sparse Mixture-of-Experts models. Listeners interested in the theory-first rigor behind cost-cutting techniques in LLM post-training will find the paper's formal proof-before-benchmarks approach a refreshing departure from the field's norm. Sources: 1. Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation — Yecheng Wu, Song Han, Hai Cai, 2026 http://arxiv.org/abs/2604.13010 2. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, Olivier Bachem (Google DeepMind), 2024 (ICLR) https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes 3. On-Policy Distillation — Kevin Lu et al. (Thinking Machines Lab), 2025 https://scholar.google.com/scholar?q=On-Policy+Distillation 4. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015 https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network 5. Tulu 3: Pushing Frontiers in Open Language Model Post-Training — Nathan Lambert, Jacob Morrison, Valentina Pyatkin, et al. (Allen Institute for AI), 2024 https://scholar.google.com/scholar?q=Tulu+3%3A+Pushing+Frontiers+in+Open+Language+Model+Post-Training 6. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI (Daya Guo et al.), 2025 https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning 7. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models — Zhihong Shao, Peiyi Wang, Qihao Zhu, et al. (DeepSeek-AI), 2024 https://scholar.google.com/scholar?q=DeepSeekMath%3A+Pushing+the+Limits+of+Mathematical+Reasoning+in+Open+Language+Models 8. Let's Verify Step by Step — Hunter Lightman, Vineet Kosaraju, Yura Burda, et al. (OpenAI), 2023 https://scholar.google.com/scholar?q=Let%27s+Verify+Step+by+Step 9. On-policy distillation (Thinking Machines Lab: Connectionism) — Kevin Lu, Thinking Machines Lab, 2025 https://scholar.google.com/scholar?q=On-policy+distillation+%28Thinking+Machines+Lab%3A+Connectionism%29 10. Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? — Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Shiji Song, Gao Huang, 2025 https://scholar.google.com/scholar?q=Does+reinforcement+learning+really+incentivize+reasoning+capacity+in+LLMs+beyond+the+base+model%3F 11. RL's razor: Why online reinforcement learning forgets less — Idan Shenfeld, Jyothish Pari, Pulkit Agrawal, 2025 https://scholar.google.com/scholar?q=RL%27s+razor%3A+Why+online+reinforcement+learning+forgets+less 12. Rethinking on-policy distillation of large language models: Phenomenology, mechanism, and recipe — Siyuan Chen, Daya Guo, Yichun Tan, Xiaohan Liang, Hao Zhou, Bo Zheng, Dejian Yang, 2026 https://scholar.google.com/scholar?q=Rethinking+on-policy+distillation+of+large+language+models%3A+Phenomenology%2C+mechanism%2C+and+recipe 13. Learning beyond teacher: Generalized on-policy distillation with reward extrapolation (ExOPD) — Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang, Saiyong Yang, Yankai Lin, 2026 https://scholar.google.com/scholar?q=Learning+beyond+teacher%3A+Generalized+on-policy+distillation+with+reward+extrapolation+%28ExOPD%29 Interactive Visualization: Lightning OPD: Precomputing the Teacher for Cheaper Reasoning Distillation

Episode metadata supplied by the publisher feed · Published Jul 21, 2026

Embed this episode

NOW PLAYING

Lightning OPD: Precomputing the Teacher for Cheaper Reasoning Distillation

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on July 21, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!