CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization episode artwork

EPISODE · Jul 31, 2026 · 20 MIN

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 78 | cs.AI Authors: Bo-Wen Zhang, Junwei He, Wen Wang, Song-Lin Lv, Wentao Ma, Rongyi Lin, Shuhan Zhong, Lan-Zhe Guo Title: CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Arxiv: http://arxiv.org/abs/2607.25659v1 Abstract: Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated tokens. This leaves no explicit mechanism for allocating credit within a response, even when different criteria are grounded in different spans, formatting decisions, or semantic choices. We propose CoRT, a token-level credit weighting method for rubric-conditioned GRPO. Instead of training an auxiliary token scoring model, CoRT uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt. The resulting tokenwise log-likelihood contrasts serve as a proxy for dependence on the rubric context. CoRT maps these contrasts to bounded, response-normalized weights and uses them to redistribute the signed GRPO advantage across tokens, without introducing an auxiliary scorer or changing the response-level reward. Experiments across instruction-tuned models and reward granularities show that CoRT improves over matched response-level GRPO in the vast majority of comparisons, with an average gain of 4.4 percentage points. The method remains competitive with learned token-level credit baselines while avoiding a separate relevance-learning stage. These results suggest that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation while retaining the simplicity and stability of GRPO.

Episode metadata supplied by the publisher feed · Published Jul 31, 2026

Embed this episode

NOW PLAYING

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

0:00 20:06

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 20 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on July 31, 2026.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!