LLM-grounded Video Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lian, Long, Shi, Baifeng, Yala, Adam, Darrell, Trevor, Li, Boyi
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916234760552448
author Lian, Long
Shi, Baifeng
Yala, Adam
Darrell, Trevor
Li, Boyi
author_facet Lian, Long
Shi, Baifeng
Yala, Adam
Darrell, Trevor
Li, Boyi
contents Text-conditioned diffusion models have emerged as a promising tool for neural video generation. However, current models still struggle with intricate spatiotemporal prompts and often generate restricted or incorrect motion. To address these limitations, we introduce LLM-grounded Video Diffusion (LVD). Instead of directly generating videos from the text inputs, LVD first leverages a large language model (LLM) to generate dynamic scene layouts based on the text inputs and subsequently uses the generated layouts to guide a diffusion model for video generation. We show that LLMs are able to understand complex spatiotemporal dynamics from text alone and generate layouts that align closely with both the prompts and the object motion patterns typically observed in the real world. We then propose to guide video diffusion models with these layouts by adjusting the attention maps. Our approach is training-free and can be integrated into any video diffusion model that admits classifier guidance. Our results demonstrate that LVD significantly outperforms its base video diffusion model and several strong baseline methods in faithfully generating videos with the desired attributes and motion patterns.
format Preprint
id arxiv_https___arxiv_org_abs_2309_17444
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LLM-grounded Video Diffusion Models
Lian, Long
Shi, Baifeng
Yala, Adam
Darrell, Trevor
Li, Boyi
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Text-conditioned diffusion models have emerged as a promising tool for neural video generation. However, current models still struggle with intricate spatiotemporal prompts and often generate restricted or incorrect motion. To address these limitations, we introduce LLM-grounded Video Diffusion (LVD). Instead of directly generating videos from the text inputs, LVD first leverages a large language model (LLM) to generate dynamic scene layouts based on the text inputs and subsequently uses the generated layouts to guide a diffusion model for video generation. We show that LLMs are able to understand complex spatiotemporal dynamics from text alone and generate layouts that align closely with both the prompts and the object motion patterns typically observed in the real world. We then propose to guide video diffusion models with these layouts by adjusting the attention maps. Our approach is training-free and can be integrated into any video diffusion model that admits classifier guidance. Our results demonstrate that LVD significantly outperforms its base video diffusion model and several strong baseline methods in faithfully generating videos with the desired attributes and motion patterns.
title LLM-grounded Video Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2309.17444