End-to-End Training for Autoregressive Video Diffusion via Self-Resampling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guo, Yuwei, Yang, Ceyuan, He, Hao, Zhao, Yang, Wei, Meng, Yang, Zhenheng, Huang, Weilin, Lin, Dahua
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912771337093120
author Guo, Yuwei
Yang, Ceyuan
He, Hao
Zhao, Yang
Wei, Meng
Yang, Zhenheng
Huang, Weilin
Lin, Dahua
author_facet Guo, Yuwei
Yang, Ceyuan
He, Hao
Zhao, Yang
Wei, Meng
Yang, Zhenheng
Huang, Weilin
Lin, Dahua
contents Autoregressive video diffusion models hold promise for world simulation but are vulnerable to exposure bias arising from the train-test mismatch. While recent works address this via post-training, they typically rely on a bidirectional teacher model or online discriminator. To achieve an end-to-end solution, we introduce Resampling Forcing, a teacher-free framework that enables training autoregressive video models from scratch and at scale. Central to our approach is a self-resampling scheme that simulates inference-time model errors on history frames during training. Conditioned on these degraded histories, a sparse causal mask enforces temporal causality while enabling parallel training with frame-level diffusion loss. To facilitate efficient long-horizon generation, we further introduce history routing, a parameter-free mechanism that dynamically retrieves the top-k most relevant history frames for each query. Experiments demonstrate that our approach achieves performance comparable to distillation-based baselines while exhibiting superior temporal consistency on longer videos owing to native-length training.
format Preprint
id arxiv_https___arxiv_org_abs_2512_15702
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
Guo, Yuwei
Yang, Ceyuan
He, Hao
Zhao, Yang
Wei, Meng
Yang, Zhenheng
Huang, Weilin
Lin, Dahua
Computer Vision and Pattern Recognition
Autoregressive video diffusion models hold promise for world simulation but are vulnerable to exposure bias arising from the train-test mismatch. While recent works address this via post-training, they typically rely on a bidirectional teacher model or online discriminator. To achieve an end-to-end solution, we introduce Resampling Forcing, a teacher-free framework that enables training autoregressive video models from scratch and at scale. Central to our approach is a self-resampling scheme that simulates inference-time model errors on history frames during training. Conditioned on these degraded histories, a sparse causal mask enforces temporal causality while enabling parallel training with frame-level diffusion loss. To facilitate efficient long-horizon generation, we further introduce history routing, a parameter-free mechanism that dynamically retrieves the top-k most relevant history frames for each query. Experiments demonstrate that our approach achieves performance comparable to distillation-based baselines while exhibiting superior temporal consistency on longer videos owing to native-length training.
title End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.15702