Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement episode artwork

EPISODE · Sep 1, 2026 · 22 MIN

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 36 | cs.LG, cs.CL Authors: Yi Ding, Ruqi Zhang Title: Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Arxiv: http://arxiv.org/abs/2608.31046v1 Abstract: On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR). However, the teacher scores student-generated trajectories that are inherently off-policy for it, so the reliability of its supervision, and hence the source of the student's improvement, remains unclear. We quantitatively analyze teacher supervision during OPD training and find substantial noise whose prevalence increases with teacher scale. Surprisingly, the student policy is insensitive to such noise, converging to comparable performance regardless of whether noisy supervision is retained or removed. Does OPD distill at all? By analyzing what drives its gains, we find that learning concentrates on low log-probability tokens, and using a single fixed negative advantage matches the performance of teacher-provided ones. This suggests that OPD works largely by suppressing low log-probability tokens, which requires no teacher. These findings motivate On-Policy Self-Adaptation (OPSA), a supervision-free method using entropy-adaptive negative advantages. It assigns stronger learning signals to high-entropy positions, suppressing tail tokens, and evenly redistributing probability mass among head tokens. Compared with the base \texttt{Qwen3-1.7B}, OPSA improves Avg@32 by 35.41 points on AIME24, corresponding to a 263\% relative gain, and more than doubles Pass@32 across all three benchmarks. It also outperforms OPD by 16.77 points in Avg@32 on AIME24. Extensive experiments and analyses across model families and tasks further demonstrate its effectiveness and generalizability.

Episode metadata supplied by the publisher feed · Published Sep 1, 2026

Embed this episode

NOW PLAYING

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

0:00 22:04

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 22 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on September 1, 2026.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!