FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiu, Haonan, Xia, Menghan, Zhang, Yong, He, Yingqing, Wang, Xintao, Shan, Ying, Liu, Ziwei
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910310694125568
author Qiu, Haonan
Xia, Menghan
Zhang, Yong
He, Yingqing
Wang, Xintao
Shan, Ying
Liu, Ziwei
author_facet Qiu, Haonan
Xia, Menghan
Zhang, Yong
He, Yingqing
Wang, Xintao
Shan, Ying
Liu, Ziwei
contents With the availability of large-scale video datasets and the advances of diffusion models, text-driven video generation has achieved substantial progress. However, existing video generation models are typically trained on a limited number of frames, resulting in the inability to generate high-fidelity long videos during inference. Furthermore, these models only support single-text conditions, whereas real-life scenarios often require multi-text conditions as the video content changes over time. To tackle these challenges, this study explores the potential of extending the text-driven capability to generate longer videos conditioned on multiple texts. 1) We first analyze the impact of initial noise in video diffusion models. Then building upon the observation of noise, we propose FreeNoise, a tuning-free and time-efficient paradigm to enhance the generative capabilities of pretrained video diffusion models while preserving content consistency. Specifically, instead of initializing noises for all frames, we reschedule a sequence of noises for long-range correlation and perform temporal attention over them by window-based function. 2) Additionally, we design a novel motion injection method to support the generation of videos conditioned on multiple text prompts. Extensive experiments validate the superiority of our paradigm in extending the generative capabilities of video diffusion models. It is noteworthy that compared with the previous best-performing method which brought about 255% extra time cost, our method incurs only negligible time cost of approximately 17%. Generated video samples are available at our website: http://haonanqiu.com/projects/FreeNoise.html.
format Preprint
id arxiv_https___arxiv_org_abs_2310_15169
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling
Qiu, Haonan
Xia, Menghan
Zhang, Yong
He, Yingqing
Wang, Xintao
Shan, Ying
Liu, Ziwei
Computer Vision and Pattern Recognition
With the availability of large-scale video datasets and the advances of diffusion models, text-driven video generation has achieved substantial progress. However, existing video generation models are typically trained on a limited number of frames, resulting in the inability to generate high-fidelity long videos during inference. Furthermore, these models only support single-text conditions, whereas real-life scenarios often require multi-text conditions as the video content changes over time. To tackle these challenges, this study explores the potential of extending the text-driven capability to generate longer videos conditioned on multiple texts. 1) We first analyze the impact of initial noise in video diffusion models. Then building upon the observation of noise, we propose FreeNoise, a tuning-free and time-efficient paradigm to enhance the generative capabilities of pretrained video diffusion models while preserving content consistency. Specifically, instead of initializing noises for all frames, we reschedule a sequence of noises for long-range correlation and perform temporal attention over them by window-based function. 2) Additionally, we design a novel motion injection method to support the generation of videos conditioned on multiple text prompts. Extensive experiments validate the superiority of our paradigm in extending the generative capabilities of video diffusion models. It is noteworthy that compared with the previous best-performing method which brought about 255% extra time cost, our method incurs only negligible time cost of approximately 17%. Generated video samples are available at our website: http://haonanqiu.com/projects/FreeNoise.html.
title FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2310.15169