Astrolabe: Steering Forward-Process Reinforcement Learning for Distilled Autoregressive Video Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Songchun, Xue, Zeyue, Fu, Siming, Huang, Jie, Kong, Xianghao, Ma, Y, Huang, Haoyang, Duan, Nan, Rao, Anyi
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915870977032192
author Zhang, Songchun
Xue, Zeyue
Fu, Siming
Huang, Jie
Kong, Xianghao
Ma, Y
Huang, Haoyang
Duan, Nan
Rao, Anyi
author_facet Zhang, Songchun
Xue, Zeyue
Fu, Siming
Huang, Jie
Kong, Xianghao
Ma, Y
Huang, Haoyang
Duan, Nan
Rao, Anyi
contents Distilled autoregressive (AR) video models enable efficient streaming generation but frequently misalign with human visual preferences. Existing reinforcement learning (RL) frameworks are not naturally suited to these architectures, typically requiring either expensive re-distillation or solver-coupled reverse-process optimization that introduces considerable memory and computational overhead. We present Astrolabe, an efficient online RL framework tailored for distilled AR models. To overcome existing bottlenecks, we introduce a forward-process RL formulation based on negative-aware fine-tuning. By contrasting positive and negative samples directly at inference endpoints, this approach establishes an implicit policy improvement direction without requiring reverse-process unrolling. To scale this alignment to long videos, we propose a streaming training scheme that generates sequences progressively via a rolling KV-cache, applying RL updates exclusively to local clip windows while conditioning on prior context to ensure long-range coherence. Finally, to mitigate reward hacking, we integrate a multi-reward objective stabilized by uncertainty-aware selective regularization and dynamic reference updates. Extensive experiments demonstrate that our method consistently enhances generation quality across multiple distilled AR video models, serving as a robust and scalable alignment solution.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17051
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Astrolabe: Steering Forward-Process Reinforcement Learning for Distilled Autoregressive Video Models
Zhang, Songchun
Xue, Zeyue
Fu, Siming
Huang, Jie
Kong, Xianghao
Ma, Y
Huang, Haoyang
Duan, Nan
Rao, Anyi
Computer Vision and Pattern Recognition
Distilled autoregressive (AR) video models enable efficient streaming generation but frequently misalign with human visual preferences. Existing reinforcement learning (RL) frameworks are not naturally suited to these architectures, typically requiring either expensive re-distillation or solver-coupled reverse-process optimization that introduces considerable memory and computational overhead. We present Astrolabe, an efficient online RL framework tailored for distilled AR models. To overcome existing bottlenecks, we introduce a forward-process RL formulation based on negative-aware fine-tuning. By contrasting positive and negative samples directly at inference endpoints, this approach establishes an implicit policy improvement direction without requiring reverse-process unrolling. To scale this alignment to long videos, we propose a streaming training scheme that generates sequences progressively via a rolling KV-cache, applying RL updates exclusively to local clip windows while conditioning on prior context to ensure long-range coherence. Finally, to mitigate reward hacking, we integrate a multi-reward objective stabilized by uncertainty-aware selective regularization and dynamic reference updates. Extensive experiments demonstrate that our method consistently enhances generation quality across multiple distilled AR video models, serving as a robust and scalable alignment solution.
title Astrolabe: Steering Forward-Process Reinforcement Learning for Distilled Autoregressive Video Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.17051