VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models episode artwork

EPISODE · Sep 27, 2025 · 22 MIN

VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 95 | cs.LG, cs.CL Authors: Guochao Jiang, Wenfeng Feng, Guofeng Quan, Chuzhan Hao, Yuewei Zhang, Guohua Liu, Hao Wang Title: VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models Arxiv: http://arxiv.org/abs/2509.19803v1 Abstract: Policy-based reinforcement learning currently plays an important role in improving LLMs on mathematical reasoning tasks. However, existing rollout-based reinforcement learning methods (GRPO, DAPO, GSPO, etc.) fail to explicitly consider LLMs' learning ability for samples of different difficulty levels, which is contrary to the human cognitive process of mathematical reasoning tasks from easy to difficult. Intuitively, we find that the variance of the rollout group's reward in RLVR partly reflects the difficulty of the current sample for LLMs. Samples that are too easy or too difficult have a lower variance, while samples with moderate difficulty have a higher variance. Based on this, we propose VCRL, a curriculum reinforcement learning framework that dynamically controls the difficulty of training samples based on the variance of group rewards. Experiments on five mathematical benchmarks and two models reveal the advantages of VCRL over the current LLM RL baselines.

Episode metadata supplied by the publisher feed · Published Sep 27, 2025

🤗 Upvotes: 95 | cs.LG, cs.CL Authors: Guochao Jiang, Wenfeng Feng, Guofeng Quan, Chuzhan Hao, Yuewei Zhang, Guohua Liu, Hao Wang Title: VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models Arxiv: http://arxiv.org/abs/2509.19803v1 Abstract: Policy-based reinforcement learning currently plays an important role in improving LLMs on mathematical reasoning tasks. However, existing rollout-based reinforcement learning methods (GRPO, DAPO, GSPO, etc.) fail to explicitly consider LLMs' learning ability for samples of different difficulty levels, which is contrary to the human cognitive process of mathematical reasoning tasks from easy to difficult. Intuitively, we find that the variance of the rollout group's reward in RLVR partly reflects the difficulty of the current sample for LLMs. Samples that are too easy or too difficult have a lower variance, while samples with moderate difficulty have a higher variance. Based on this, we propose VCRL, a curriculum reinforcement learning framework that dynamically controls the difficulty of training samples based on the variance of group rewards. Experiments on five mathematical benchmarks and two models reveal the advantages of VCRL over the current LLM RL baselines.

PodParley-generated summary based on available episode metadata and transcript content.

NOW PLAYING

VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models

0:00 22:17

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 22 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on September 27, 2025.

What is this episode about?

🤗 Upvotes: 95 | cs.LG, cs.CL Authors: Guochao Jiang, Wenfeng Feng, Guofeng Quan, Chuzhan Hao, Yuewei Zhang, Guohua Liu, Hao Wang Title: VCRL: Variance-based Curriculum Reinforcement...

Can I download this Daily Paper Cast episode?

Yes, you can download this episode by clicking the download button on the episode player, or subscribe to the podcast in your preferred podcast app for automatic downloads.
URL copied to clipboard!