Dual-Stream Diffusion Net for Text-to-Video Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Binhui, Liu, Xin, Dai, Anbo, Zeng, Zhiyong, Wang, Dan, Cui, Zhen, Yang, Jian
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910284265816064
author Liu, Binhui
Liu, Xin
Dai, Anbo
Zeng, Zhiyong
Wang, Dan
Cui, Zhen
Yang, Jian
author_facet Liu, Binhui
Liu, Xin
Dai, Anbo
Zeng, Zhiyong
Wang, Dan
Cui, Zhen
Yang, Jian
contents With the emerging diffusion models, recently, text-to-video generation has aroused increasing attention. But an important bottleneck therein is that generative videos often tend to carry some flickers and artifacts. In this work, we propose a dual-stream diffusion net (DSDN) to improve the consistency of content variations in generating videos. In particular, the designed two diffusion streams, video content and motion branches, could not only run separately in their private spaces for producing personalized video variations as well as content, but also be well-aligned between the content and motion domains through leveraging our designed cross-transformer interaction module, which would benefit the smoothness of generated videos. Besides, we also introduce motion decomposer and combiner to faciliate the operation on video motion. Qualitative and quantitative experiments demonstrate that our method could produce amazing continuous videos with fewer flickers.
format Preprint
id arxiv_https___arxiv_org_abs_2308_08316
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Dual-Stream Diffusion Net for Text-to-Video Generation
Liu, Binhui
Liu, Xin
Dai, Anbo
Zeng, Zhiyong
Wang, Dan
Cui, Zhen
Yang, Jian
Computer Vision and Pattern Recognition
With the emerging diffusion models, recently, text-to-video generation has aroused increasing attention. But an important bottleneck therein is that generative videos often tend to carry some flickers and artifacts. In this work, we propose a dual-stream diffusion net (DSDN) to improve the consistency of content variations in generating videos. In particular, the designed two diffusion streams, video content and motion branches, could not only run separately in their private spaces for producing personalized video variations as well as content, but also be well-aligned between the content and motion domains through leveraging our designed cross-transformer interaction module, which would benefit the smoothness of generated videos. Besides, we also introduce motion decomposer and combiner to faciliate the operation on video motion. Qualitative and quantitative experiments demonstrate that our method could produce amazing continuous videos with fewer flickers.
title Dual-Stream Diffusion Net for Text-to-Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2308.08316