Saved in:
Bibliographic Details
Main Authors: Yang, Yongqi, Huang, Huayang, Peng, Xu, Hu, Xiaobin, Luo, Donghao, Zhang, Jiangning, Wang, Chengjie, Wu, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.01419
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908895334629376
author Yang, Yongqi
Huang, Huayang
Peng, Xu
Hu, Xiaobin
Luo, Donghao
Zhang, Jiangning
Wang, Chengjie
Wu, Yu
author_facet Yang, Yongqi
Huang, Huayang
Peng, Xu
Hu, Xiaobin
Luo, Donghao
Zhang, Jiangning
Wang, Chengjie
Wu, Yu
contents Recent hybrid video generation models combine autoregressive temporal dynamics with diffusion-based spatial denoising, but their sequential, iterative nature leads to error accumulation and long inference times. In this work, we propose a distillation-based framework for efficient causal video generation that enables high-quality synthesis with extremely limited denoising steps. Our approach builds upon the Distribution Matching Distillation (DMD) framework and proposes a novel Adversarial Self-Distillation (ASD) strategy, which aligns the outputs of the student model's n-step denoising process with its (n+1)-step version at the distribution level. This design provides smoother supervision by bridging small intra-student gaps and more informative guidance by combining teacher knowledge with locally consistent student behavior, substantially improving training stability and generation quality in extremely few-step scenarios (e.g., 1-2 steps). In addition, we present a First-Frame Enhancement (FFE) strategy, which allocates more denoising steps to the initial frames to mitigate error propagation while applying larger skipping steps to later frames. Extensive experiments on VBench demonstrate that our method surpasses state-of-the-art approaches in both one-step and two-step video generation. Notably, our framework produces a single distilled model that flexibly supports multiple inference-step settings, eliminating the need for repeated re-distillation and enabling efficient, high-quality video synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2511_01419
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards One-step Causal Video Generation via Adversarial Self-Distillation
Yang, Yongqi
Huang, Huayang
Peng, Xu
Hu, Xiaobin
Luo, Donghao
Zhang, Jiangning
Wang, Chengjie
Wu, Yu
Computer Vision and Pattern Recognition
Recent hybrid video generation models combine autoregressive temporal dynamics with diffusion-based spatial denoising, but their sequential, iterative nature leads to error accumulation and long inference times. In this work, we propose a distillation-based framework for efficient causal video generation that enables high-quality synthesis with extremely limited denoising steps. Our approach builds upon the Distribution Matching Distillation (DMD) framework and proposes a novel Adversarial Self-Distillation (ASD) strategy, which aligns the outputs of the student model's n-step denoising process with its (n+1)-step version at the distribution level. This design provides smoother supervision by bridging small intra-student gaps and more informative guidance by combining teacher knowledge with locally consistent student behavior, substantially improving training stability and generation quality in extremely few-step scenarios (e.g., 1-2 steps). In addition, we present a First-Frame Enhancement (FFE) strategy, which allocates more denoising steps to the initial frames to mitigate error propagation while applying larger skipping steps to later frames. Extensive experiments on VBench demonstrate that our method surpasses state-of-the-art approaches in both one-step and two-step video generation. Notably, our framework produces a single distilled model that flexibly supports multiple inference-step settings, eliminating the need for repeated re-distillation and enabling efficient, high-quality video synthesis.
title Towards One-step Causal Video Generation via Adversarial Self-Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.01419