KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration episode artwork

EPISODE · May 20, 2026 · 23 MIN

KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 36 | cs.CV Authors: Ruicheng Zhang, Kaixi Cong, Jun Zhou, Zhizhou Zhong, Zunnan Xu, Shuiyang Mao, Wei Liu, Xiu Li Title: KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration Arxiv: http://arxiv.org/abs/2605.14278v1 Abstract: Aligning streaming autoregressive (AR) video generators with human preferences is challenging. Existing reinforcement learning methods predominantly rely on noise-based exploration and SDE-based surrogate policies that are mismatched to the deterministic ODE dynamics of distilled AR models, and tend to perturb low-level appearance rather than the high-level semantic storyline progression critical for long-horizon coherence. To address these limitations, we present KVPO, an ODE-native online Group Relative Policy Optimization (GRPO) framework for aligning streaming video generators. For diversity exploration, KVPO introduces a causal-semantic exploration paradigm that relocates the source of variation from stochastic noise to the historical KV cache. By stochastically routing historical KV entries, it constructs semantically diverse generation branches that remain strictly on the data manifold. For policy modeling, KVPO introduces a velocity-field surrogate policy based on Trajectory Velocity Energy (TVE), which quantifies branch likelihood in flow-matching velocity space and yields a reward-weighted contrastive objective fully consistent with the native ODE formulation. Experiments on multiple distilled AR video generators demonstrate consistent gains in visual quality, motion quality, and text-video alignment across both single-prompt short-video and multi-prompt long-video settings.

Episode metadata supplied by the publisher feed · Published May 20, 2026

Embed this episode

NOW PLAYING

KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration

0:00 23:40

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 23 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on May 20, 2026.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!