Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies episode artwork

EPISODE · Dec 25, 2025 · 27 MIN

Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 49 | cs.LG, cs.AI, cs.CL Authors: Yuqiao Tan, Minzheng Wang, Shizhu He, Huanxuan Liao, Chengfeng Zhao, Qiunan Lu, Tian Liang, Jun Zhao, Kang Liu Title: Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies Arxiv: http://arxiv.org/abs/2512.19673v1 Abstract: Existing reinforcement learning (RL) approaches treat large language models (LLMs) as a single unified policy, overlooking their internal mechanisms. Understanding how policy evolves across layers and modules is therefore crucial for enabling more targeted optimization and raveling out complex reasoning mechanisms. In this paper, we decompose the language model policy by leveraging the intrinsic split of the Transformer residual stream and the equivalence between the composition of hidden states with the unembedding matrix and the resulting samplable policy. This decomposition reveals Internal Layer Policies, corresponding to contributions from individual layers, and Internal Modular Policies, which align with the self-attention and feed-forward network (FFN) components within each layer. By analyzing the entropy of internal policy, we find that: (a) Early layers keep high entropy for exploration, top layers converge to near-zero entropy for refinement, with convergence patterns varying across model series. (b) LLama's prediction space rapidly converges in the final layer, whereas Qwen-series models, especially Qwen3, exhibit a more human-like, progressively structured reasoning pattern. Motivated by these findings, we propose Bottom-up Policy Optimization (BuPO), a novel RL paradigm that directly optimizes the internal layer policy during early training. By aligning training objective at lower layer, BuPO reconstructs foundational reasoning capabilities and achieves superior performance. Extensive experiments on complex reasoning benchmarks demonstrates the effectiveness of our method. Our code is available at https://github.com/Trae1ounG/BuPO.

Episode metadata supplied by the publisher feed · Published Dec 25, 2025

Embed this episode

NOW PLAYING

Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies

0:00 27:28

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 27 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on December 25, 2025.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!