ExVideo: Extending Video Diffusion Models via Parameter-Efficient Post-Tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Duan, Zhongjie, Zhou, Wenmeng, Chen, Cen, Li, Yaliang, Qian, Weining
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910495521374208
author Duan, Zhongjie
Zhou, Wenmeng
Chen, Cen
Li, Yaliang
Qian, Weining
author_facet Duan, Zhongjie
Zhou, Wenmeng
Chen, Cen
Li, Yaliang
Qian, Weining
contents Recently, advancements in video synthesis have attracted significant attention. Video synthesis models such as AnimateDiff and Stable Video Diffusion have demonstrated the practical applicability of diffusion models in creating dynamic visual content. The emergence of SORA has further spotlighted the potential of video generation technologies. Nonetheless, the extension of video lengths has been constrained by the limitations in computational resources. Most existing video synthesis models can only generate short video clips. In this paper, we propose a novel post-tuning methodology for video synthesis models, called ExVideo. This approach is designed to enhance the capability of current video synthesis models, allowing them to produce content over extended temporal durations while incurring lower training expenditures. In particular, we design extension strategies across common temporal model architectures respectively, including 3D convolution, temporal attention, and positional embedding. To evaluate the efficacy of our proposed post-tuning approach, we conduct extension training on the Stable Video Diffusion model. Our approach augments the model's capacity to generate up to $5\times$ its original number of frames, requiring only 1.5k GPU hours of training on a dataset comprising 40k videos. Importantly, the substantial increase in video length doesn't compromise the model's innate generalization capabilities, and the model showcases its advantages in generating videos of diverse styles and resolutions. We will release the source code and the enhanced model publicly.
format Preprint
id arxiv_https___arxiv_org_abs_2406_14130
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ExVideo: Extending Video Diffusion Models via Parameter-Efficient Post-Tuning
Duan, Zhongjie
Zhou, Wenmeng
Chen, Cen
Li, Yaliang
Qian, Weining
Computer Vision and Pattern Recognition
Recently, advancements in video synthesis have attracted significant attention. Video synthesis models such as AnimateDiff and Stable Video Diffusion have demonstrated the practical applicability of diffusion models in creating dynamic visual content. The emergence of SORA has further spotlighted the potential of video generation technologies. Nonetheless, the extension of video lengths has been constrained by the limitations in computational resources. Most existing video synthesis models can only generate short video clips. In this paper, we propose a novel post-tuning methodology for video synthesis models, called ExVideo. This approach is designed to enhance the capability of current video synthesis models, allowing them to produce content over extended temporal durations while incurring lower training expenditures. In particular, we design extension strategies across common temporal model architectures respectively, including 3D convolution, temporal attention, and positional embedding. To evaluate the efficacy of our proposed post-tuning approach, we conduct extension training on the Stable Video Diffusion model. Our approach augments the model's capacity to generate up to $5\times$ its original number of frames, requiring only 1.5k GPU hours of training on a dataset comprising 40k videos. Importantly, the substantial increase in video length doesn't compromise the model's innate generalization capabilities, and the model showcases its advantages in generating videos of diverse styles and resolutions. We will release the source code and the enhanced model publicly.
title ExVideo: Extending Video Diffusion Models via Parameter-Efficient Post-Tuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.14130