Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Xun, Li, Zhengqi, He, Guande, Zhou, Mingyuan, Shechtman, Eli |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Causality in Video Diffusers is Separable from Denoising
by: Bai, Xingjian, et al.
Published: (2026)
by: Bai, Xingjian, et al.
Published: (2026)
Self-Evaluation Unlocks Any-Step Text-to-Image Generation
by: Yu, Xin, et al.
Published: (2025)
by: Yu, Xin, et al.
Published: (2025)
MotionStream: Real-Time Video Generation with Interactive Motion Controls
by: Shin, Joonghyuk, et al.
Published: (2025)
by: Shin, Joonghyuk, et al.
Published: (2025)
Score Distillation of Flow Matching Models
by: Zhou, Mingyuan, et al.
Published: (2025)
by: Zhou, Mingyuan, et al.
Published: (2025)
Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
by: Lin, Shanchuan, et al.
Published: (2025)
by: Lin, Shanchuan, et al.
Published: (2025)
End-to-End Training for Unified Tokenization and Latent Denoising
by: Duggal, Shivam, et al.
Published: (2026)
by: Duggal, Shivam, et al.
Published: (2026)
Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models
by: Kwon, Taesung, et al.
Published: (2026)
by: Kwon, Taesung, et al.
Published: (2026)
Layer- and Timestep-Adaptive Differentiable Token Compression Ratios for Efficient Diffusion Transformers
by: You, Haoran, et al.
Published: (2024)
by: You, Haoran, et al.
Published: (2024)
Score identity Distillation: Exponentially Fast Distillation of Pretrained Diffusion Models for One-Step Generation
by: Zhou, Mingyuan, et al.
Published: (2024)
by: Zhou, Mingyuan, et al.
Published: (2024)
Improved Baselines with Representation Autoencoders
by: Singh, Jaskirat, et al.
Published: (2026)
by: Singh, Jaskirat, et al.
Published: (2026)
Consistency Diffusion Bridge Models
by: He, Guande, et al.
Published: (2024)
by: He, Guande, et al.
Published: (2024)
REG: Rectified Gradient Guidance for Conditional Diffusion Models
by: Gao, Zhengqi, et al.
Published: (2025)
by: Gao, Zhengqi, et al.
Published: (2025)
Bridging Diversity and Uncertainty in Active learning with Self-Supervised Pre-Training
by: Doucet, Paul, et al.
Published: (2024)
by: Doucet, Paul, et al.
Published: (2024)
What matters for Representation Alignment: Global Information or Spatial Structure?
by: Singh, Jaskirat, et al.
Published: (2025)
by: Singh, Jaskirat, et al.
Published: (2025)
Control-Augmented Autoregressive Diffusion for Data Assimilation
by: Srivastava, Prakhar, et al.
Published: (2025)
by: Srivastava, Prakhar, et al.
Published: (2025)
DiNO-Diffusion. Scaling Medical Diffusion via Self-Supervised Pre-Training
by: Jimenez-Perez, Guillermo, et al.
Published: (2024)
by: Jimenez-Perez, Guillermo, et al.
Published: (2024)
Diffusion Adversarial Post-Training for One-Step Video Generation
by: Lin, Shanchuan, et al.
Published: (2025)
by: Lin, Shanchuan, et al.
Published: (2025)
BridgeDrive: Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous Driving
by: Liu, Shu, et al.
Published: (2025)
by: Liu, Shu, et al.
Published: (2025)
Bridging the Gap Between Multimodal Foundation Models and World Models
by: He, Xuehai
Published: (2025)
by: He, Xuehai
Published: (2025)
Training Video Foundation Models with NVIDIA NeMo
by: Patel, Zeeshan, et al.
Published: (2025)
by: Patel, Zeeshan, et al.
Published: (2025)
SpectralAR: Spectral Autoregressive Visual Generation
by: Huang, Yuanhui, et al.
Published: (2025)
by: Huang, Yuanhui, et al.
Published: (2025)
VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's Guide
by: Lee, Dohun, et al.
Published: (2024)
by: Lee, Dohun, et al.
Published: (2024)
CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
by: Li, Zian, et al.
Published: (2025)
by: Li, Zian, et al.
Published: (2025)
Relational Visual Similarity
by: Nguyen, Thao, et al.
Published: (2025)
by: Nguyen, Thao, et al.
Published: (2025)
FEDEXCHANGE: Bridging the Domain Gap in Federated Object Detection for Free
by: Yuan, Haolin, et al.
Published: (2025)
by: Yuan, Haolin, et al.
Published: (2025)
From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
by: Yin, Tianwei, et al.
Published: (2024)
by: Yin, Tianwei, et al.
Published: (2024)
FrameBridge: Improving Image-to-Video Generation with Bridge Models
by: Wang, Yuji, et al.
Published: (2024)
by: Wang, Yuji, et al.
Published: (2024)
Denoising Score Distillation: From Noisy Diffusion Pretraining to One-Step High-Quality Generation
by: Chen, Tianyu, et al.
Published: (2025)
by: Chen, Tianyu, et al.
Published: (2025)
Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
by: Lokegaonkar, Vaibhavi, et al.
Published: (2026)
by: Lokegaonkar, Vaibhavi, et al.
Published: (2026)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Rethinking Training Dynamics in Scale-wise Autoregressive Generation
by: Zhou, Gengze, et al.
Published: (2025)
by: Zhou, Gengze, et al.
Published: (2025)
Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts
by: Yuan, Yike, et al.
Published: (2025)
by: Yuan, Yike, et al.
Published: (2025)
Diffusion Beats Autoregressive in Data-Constrained Settings
by: Prabhudesai, Mihir, et al.
Published: (2025)
by: Prabhudesai, Mihir, et al.
Published: (2025)
When Test-Time Guidance Is Enough: Fast Image and Video Editing with Diffusion Guidance
by: Ghorbel, Ahmed, et al.
Published: (2026)
by: Ghorbel, Ahmed, et al.
Published: (2026)
VideoPDE: Unified Generative PDE Solving via Video Inpainting Diffusion Models
by: Li, Edward, et al.
Published: (2025)
by: Li, Edward, et al.
Published: (2025)
Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models
by: Chen, Zijun, et al.
Published: (2024)
by: Chen, Zijun, et al.
Published: (2024)
Pixel-Space Post-Training of Latent Diffusion Models
by: Zhang, Christina, et al.
Published: (2024)
by: Zhang, Christina, et al.
Published: (2024)
Scaling Image and Video Generation via Test-Time Evolutionary Search
by: He, Haoran, et al.
Published: (2025)
by: He, Haoran, et al.
Published: (2025)
MASC: Boosting Autoregressive Image Generation with a Manifold-Aligned Semantic Clustering
by: He, Lixuan, et al.
Published: (2025)
by: He, Lixuan, et al.
Published: (2025)
DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhancer
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
Similar Items
-
Causality in Video Diffusers is Separable from Denoising
by: Bai, Xingjian, et al.
Published: (2026) -
Self-Evaluation Unlocks Any-Step Text-to-Image Generation
by: Yu, Xin, et al.
Published: (2025) -
MotionStream: Real-Time Video Generation with Interactive Motion Controls
by: Shin, Joonghyuk, et al.
Published: (2025) -
Score Distillation of Flow Matching Models
by: Zhou, Mingyuan, et al.
Published: (2025) -
Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
by: Lin, Shanchuan, et al.
Published: (2025)