Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe episode artwork

EPISODE · Apr 16, 2026 · 25 MIN

Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 62 | cs.LG, cs.AI, cs.CL Authors: Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, Tianyu Yu, Huan-ang Gao, Wenkai Yang, Zhiyuan Liu, Ning Ding Title: Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe Arxiv: http://arxiv.org/abs/2604.13016v2 Abstract: On-policy distillation (OPD) has become a core technique in the post-training of large language models, yet its training dynamics remain poorly understood. This paper provides a systematic investigation of OPD dynamics and mechanisms. We first identify that two conditions govern whether OPD succeeds or fails: (i) the student and teacher should share compatible thinking patterns; and (ii) even with consistent thinking patterns and higher scores, the teacher must offer genuinely new capabilities beyond what the student has seen during training. We validate these findings through weak-to-strong reverse distillation, showing that same-family 1.5B and 7B teachers are distributionally indistinguishable from the student's perspective. Probing into the token-level mechanism, we show that successful OPD is characterized by progressive alignment on high-probability tokens at student-visited states, a small shared token set that concentrates most of the probability mass (97%-99%). We further propose two practical strategies to recover failing OPD: off-policy cold start and teacher-aligned prompt selection. Finally, we show that OPD's apparent free lunch of dense token-level reward comes at a cost, raising the question of whether OPD can scale to long-horizon distillation.

Episode metadata supplied by the publisher feed · Published Apr 16, 2026

Embed this episode

NOW PLAYING

Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

0:00 25:07

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 25 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on April 16, 2026.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!