AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Haiyu, Chen, Xinyuan, Wang, Yaohui, Liu, Xihui, Wang, Yunhong, Qiao, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910892507004928
author Zhang, Haiyu
Chen, Xinyuan
Wang, Yaohui
Liu, Xihui
Wang, Yunhong
Qiao, Yu
author_facet Zhang, Haiyu
Chen, Xinyuan
Wang, Yaohui
Liu, Xihui
Wang, Yunhong
Qiao, Yu
contents Diffusion models have achieved remarkable progress in the field of video generation. However, their iterative denoising nature requires a large number of inference steps to generate a video, which is slow and computationally expensive. In this paper, we begin with a detailed analysis of the challenges present in existing diffusion distillation methods and propose a novel efficient method, namely AccVideo, to reduce the inference steps for accelerating video diffusion models with synthetic dataset. We leverage the pretrained video diffusion model to generate multiple valid denoising trajectories as our synthetic dataset, which eliminates the use of useless data points during distillation. Based on the synthetic dataset, we design a trajectory-based few-step guidance that utilizes key data points from the denoising trajectories to learn the noise-to-video mapping, enabling video generation in fewer steps. Furthermore, since the synthetic dataset captures the data distribution at each diffusion timestep, we introduce an adversarial training strategy to align the output distribution of the student model with that of our synthetic dataset, thereby enhancing the video quality. Extensive experiments demonstrate that our model achieves 8.5x improvements in generation speed compared to the teacher model while maintaining comparable performance. Compared to previous accelerating methods, our approach is capable of generating videos with higher quality and resolution, i.e., 5-seconds, 720x1280, 24fps.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19462
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset
Zhang, Haiyu
Chen, Xinyuan
Wang, Yaohui
Liu, Xihui
Wang, Yunhong
Qiao, Yu
Computer Vision and Pattern Recognition
Diffusion models have achieved remarkable progress in the field of video generation. However, their iterative denoising nature requires a large number of inference steps to generate a video, which is slow and computationally expensive. In this paper, we begin with a detailed analysis of the challenges present in existing diffusion distillation methods and propose a novel efficient method, namely AccVideo, to reduce the inference steps for accelerating video diffusion models with synthetic dataset. We leverage the pretrained video diffusion model to generate multiple valid denoising trajectories as our synthetic dataset, which eliminates the use of useless data points during distillation. Based on the synthetic dataset, we design a trajectory-based few-step guidance that utilizes key data points from the denoising trajectories to learn the noise-to-video mapping, enabling video generation in fewer steps. Furthermore, since the synthetic dataset captures the data distribution at each diffusion timestep, we introduce an adversarial training strategy to align the output distribution of the student model with that of our synthetic dataset, thereby enhancing the video quality. Extensive experiments demonstrate that our model achieves 8.5x improvements in generation speed compared to the teacher model while maintaining comparable performance. Compared to previous accelerating methods, our approach is capable of generating videos with higher quality and resolution, i.e., 5-seconds, 720x1280, 24fps.
title AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.19462