Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Fei, Hao, Wu, Shengqiong, Ji, Wei, Zhang, Hanwang, Chua, Tat-Seng
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929281767047168
author Fei, Hao
Wu, Shengqiong
Ji, Wei
Zhang, Hanwang
Chua, Tat-Seng
author_facet Fei, Hao
Wu, Shengqiong
Ji, Wei
Zhang, Hanwang
Chua, Tat-Seng
contents Text-to-video (T2V) synthesis has gained increasing attention in the community, in which the recently emerged diffusion models (DMs) have promisingly shown stronger performance than the past approaches. While existing state-of-the-art DMs are competent to achieve high-resolution video generation, they may largely suffer from key limitations (e.g., action occurrence disorders, crude video motions) with respect to the intricate temporal dynamics modeling, one of the crux of video synthesis. In this work, we investigate strengthening the awareness of video dynamics for DMs, for high-quality T2V generation. Inspired by human intuition, we design an innovative dynamic scene manager (dubbed as Dysen) module, which includes (step-1) extracting from input text the key actions with proper time-order arrangement, (step-2) transforming the action schedules into the dynamic scene graph (DSG) representations, and (step-3) enriching the scenes in the DSG with sufficient and reasonable details. Taking advantage of the existing powerful LLMs (e.g., ChatGPT) via in-context learning, Dysen realizes (nearly) human-level temporal dynamics understanding. Finally, the resulting video DSG with rich action scene details is encoded as fine-grained spatio-temporal features, integrated into the backbone T2V DM for video generating. Experiments on popular T2V datasets suggest that our Dysen-VDM consistently outperforms prior arts with significant margins, especially in scenarios with complex actions. Codes at https://haofei.vip/Dysen-VDM
format Preprint
id arxiv_https___arxiv_org_abs_2308_13812
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
Fei, Hao
Wu, Shengqiong
Ji, Wei
Zhang, Hanwang
Chua, Tat-Seng
Artificial Intelligence
Computer Vision and Pattern Recognition
Text-to-video (T2V) synthesis has gained increasing attention in the community, in which the recently emerged diffusion models (DMs) have promisingly shown stronger performance than the past approaches. While existing state-of-the-art DMs are competent to achieve high-resolution video generation, they may largely suffer from key limitations (e.g., action occurrence disorders, crude video motions) with respect to the intricate temporal dynamics modeling, one of the crux of video synthesis. In this work, we investigate strengthening the awareness of video dynamics for DMs, for high-quality T2V generation. Inspired by human intuition, we design an innovative dynamic scene manager (dubbed as Dysen) module, which includes (step-1) extracting from input text the key actions with proper time-order arrangement, (step-2) transforming the action schedules into the dynamic scene graph (DSG) representations, and (step-3) enriching the scenes in the DSG with sufficient and reasonable details. Taking advantage of the existing powerful LLMs (e.g., ChatGPT) via in-context learning, Dysen realizes (nearly) human-level temporal dynamics understanding. Finally, the resulting video DSG with rich action scene details is encoded as fine-grained spatio-temporal features, integrated into the backbone T2V DM for video generating. Experiments on popular T2V datasets suggest that our Dysen-VDM consistently outperforms prior arts with significant margins, especially in scenarios with complex actions. Codes at https://haofei.vip/Dysen-VDM
title Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2308.13812