Astrolabe: Steering Forward-Process Reinforcement Learning for Distilled Autoregressive Video Models episode artwork

EPISODE · Mar 24, 2026 · 24 MIN

Astrolabe: Steering Forward-Process Reinforcement Learning for Distilled Autoregressive Video Models

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 87 | cs.CV Authors: Songchun Zhang, Zeyue Xue, Siming Fu, Jie Huang, Xianghao Kong, Y Ma, Haoyang Huang, Nan Duan, Anyi Rao Title: Astrolabe: Steering Forward-Process Reinforcement Learning for Distilled Autoregressive Video Models Arxiv: http://arxiv.org/abs/2603.17051v1 Abstract: Distilled autoregressive (AR) video models enable efficient streaming generation but frequently misalign with human visual preferences. Existing reinforcement learning (RL) frameworks are not naturally suited to these architectures, typically requiring either expensive re-distillation or solver-coupled reverse-process optimization that introduces considerable memory and computational overhead. We present Astrolabe, an efficient online RL framework tailored for distilled AR models. To overcome existing bottlenecks, we introduce a forward-process RL formulation based on negative-aware fine-tuning. By contrasting positive and negative samples directly at inference endpoints, this approach establishes an implicit policy improvement direction without requiring reverse-process unrolling. To scale this alignment to long videos, we propose a streaming training scheme that generates sequences progressively via a rolling KV-cache, applying RL updates exclusively to local clip windows while conditioning on prior context to ensure long-range coherence. Finally, to mitigate reward hacking, we integrate a multi-reward objective stabilized by uncertainty-aware selective regularization and dynamic reference updates. Extensive experiments demonstrate that our method consistently enhances generation quality across multiple distilled AR video models, serving as a robust and scalable alignment solution.

Episode metadata supplied by the publisher feed · Published Mar 24, 2026

Embed this episode

NOW PLAYING

Astrolabe: Steering Forward-Process Reinforcement Learning for Distilled Autoregressive Video Models

0:00 24:19

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 24 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on March 24, 2026.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!