From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space episode artwork

EPISODE · Apr 17, 2026 · 23 MIN

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 23 | cs.LG, cs.AI, cs.CL Authors: Yuqiao Tan, Minzheng Wang, Bo Liu, Zichen Liu, Tian Liang, Shizhu He, Jun Zhao, Kang Liu Title: From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space Arxiv: http://arxiv.org/abs/2604.14142v1 Abstract: While reinforcement learning with verifiable rewards (RLVR) significantly enhances LLM reasoning by optimizing the conditional distribution P(y|x), its potential is fundamentally bounded by the base model's existing output distribution. Optimizing the marginal distribution P(y) in the Pre-train Space addresses this bottleneck by encoding reasoning ability and preserving broad exploration capacity. Yet, conventional pre-training relies on static corpora for passive learning, leading to a distribution shift that hinders targeted reasoning enhancement. In this paper, we introduce PreRL (Pre-train Space RL), which applies reward-driven online updates directly to P(y). We theoretically and empirically validate the strong gradient alignment between log P(y) and log P(y|x), establishing PreRL as a viable surrogate for standard RL. Furthermore, we uncover a critical mechanism: Negative Sample Reinforcement (NSR) within PreRL serves as an exceptionally effective driver for reasoning. NSR-PreRL rapidly prunes incorrect reasoning spaces while stimulating endogenous reflective behaviors, increasing transition and reflection thoughts by 14.89x and 6.54x, respectively. Leveraging these insights, we propose Dual Space RL (DSRL), a Policy Reincarnation strategy that initializes models with NSR-PreRL to expand the reasoning horizon before transitioning to standard RL for fine-grained optimization. Extensive experiments demonstrate that DSRL consistently outperforms strong baselines, proving that pre-train space pruning effectively steers the policy toward a refined correct reasoning subspace.

Episode metadata supplied by the publisher feed · Published Apr 17, 2026

Embed this episode

NOW PLAYING

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space

0:00 23:31

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 23 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on April 17, 2026.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!