GVDIFF: Grounded Text-to-Video Generation with Diffusion Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dou, Huanzhang, Li, Ruixiang, Su, Wei, Li, Xi
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916311448158208
author Dou, Huanzhang
Li, Ruixiang
Su, Wei
Li, Xi
author_facet Dou, Huanzhang
Li, Ruixiang
Su, Wei
Li, Xi
contents In text-to-video (T2V) generation, significant attention has been directed toward its development, yet unifying discrete and continuous grounding conditions in T2V generation remains under-explored. This paper proposes a Grounded text-to-Video generation framework, termed GVDIFF. First, we inject the grounding condition into the self-attention through an uncertainty-based representation to explicitly guide the focus of the network. Second, we introduce a spatial-temporal grounding layer that connects the grounding condition with target objects and enables the model with the grounded generation capacity in the spatial-temporal domain. Third, our dynamic gate network adaptively skips the redundant grounding process to selectively extract grounding information and semantics while improving efficiency. We extensively evaluate the grounded generation capacity of GVDIFF and demonstrate its versatility in applications, including long-range video generation, sequential prompts, and object-specific editing.
format Preprint
id arxiv_https___arxiv_org_abs_2407_01921
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GVDIFF: Grounded Text-to-Video Generation with Diffusion Models
Dou, Huanzhang
Li, Ruixiang
Su, Wei
Li, Xi
Computer Vision and Pattern Recognition
In text-to-video (T2V) generation, significant attention has been directed toward its development, yet unifying discrete and continuous grounding conditions in T2V generation remains under-explored. This paper proposes a Grounded text-to-Video generation framework, termed GVDIFF. First, we inject the grounding condition into the self-attention through an uncertainty-based representation to explicitly guide the focus of the network. Second, we introduce a spatial-temporal grounding layer that connects the grounding condition with target objects and enables the model with the grounded generation capacity in the spatial-temporal domain. Third, our dynamic gate network adaptively skips the redundant grounding process to selectively extract grounding information and semantics while improving efficiency. We extensively evaluate the grounded generation capacity of GVDIFF and demonstrate its versatility in applications, including long-range video generation, sequential prompts, and object-specific editing.
title GVDIFF: Grounded Text-to-Video Generation with Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.01921