Decoupling KL Direction from Rollout Source in LLM Distillation episode artwork

EPISODE · Aug 28, 2026

Decoupling KL Direction from Rollout Source in LLM Distillation

from AI Post Transformers

This episode examines "Decoupling KL and Trajectories," which challenges the assumption that off-policy training must pair with forward KL divergence and on-policy training with reverse KL — a coupling used by DeepSeek-R1, Qwen3, MiMo, and GLM-5 without ever being tested. The hosts unpack the two independent design axes at play: prefix source (teacher-generated versus student-generated rollouts) and KL direction (forward's distribution-covering behavior versus reverse's mode-seeking concentration), showing that standard SFT is actually just forward-KL distillation on teacher text in disguise. The paper's real contribution is exploring two previously unstudied combinations — teacher-prefix reverse KL and student-prefix forward KL — turning an assumed two-option choice into a full four-quadrant design space. Grounding the discussion is a concrete experimental setup using Qwen3-0.6B-Base as a student distilling from Qwen3-4B and Qwen3-8B teachers. Listeners interested in LLM distillation, on-policy training, or exposure bias will find this a sharp dismantling of an industry-wide convention nobody had thought to question. Sources: 1. Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation — Anhao Zhao, Haoran Xin, Yingqi Fan, Junlong Tong, Wenjie Li, Xiaoyu Shen, 2026 http://arxiv.org/abs/2605.16826 2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning 3. Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks — Samy Bengio, Oriol Vinyals, Navdeep Jaitly, Noam Shazeer, 2015 https://scholar.google.com/scholar?q=Scheduled+Sampling+for+Sequence+Prediction+with+Recurrent+Neural+Networks 4. Sequence Level Training with Recurrent Neural Networks — Marc'Aurelio Ranzato, Sumit Chopra, Michael Auli, Wojciech Zaremba, 2016 https://scholar.google.com/scholar?q=Sequence+Level+Training+with+Recurrent+Neural+Networks 5. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, Olivier Bachem (Google DeepMind), 2024 https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes 6. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2024 https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models 7. Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models — Feng Luo, Yu-Neng Chuang, Guanchu Wang, Zicheng Xu, Xiaotian Han, Tianyi Zhang, Vladimir Braverman, 2026 https://scholar.google.com/scholar?q=Demystifying+OPD%3A+Length+Inflation+and+Stabilization+Strategies+for+Large+Language+Models 8. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey J. Gordon, J. Andrew Bagnell, 2011 https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29 9. Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting — Xinyuan Chen, Noam Razin, Karthik Narasimhan, Danqi Chen, 2025 https://scholar.google.com/scholar?q=Retaining+by+Doing%3A+The+Role+of+On-Policy+Data+in+Mitigating+Forgetting Interactive Visualization: Decoupling KL Direction from Rollout Source in LLM Distillation

Episode metadata supplied by the publisher feed · Published Aug 28, 2026

Embed this episode

NOW PLAYING

Decoupling KL Direction from Rollout Source in LLM Distillation

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on August 28, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!