Comp4D: LLM-Guided Compositional 4D Scene Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Dejia, Liang, Hanwen, Bhatt, Neel P., Hu, Hezhen, Liang, Hanxue, Plataniotis, Konstantinos N., Wang, Zhangyang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909149543006208
author Xu, Dejia
Liang, Hanwen
Bhatt, Neel P.
Hu, Hezhen
Liang, Hanxue
Plataniotis, Konstantinos N.
Wang, Zhangyang
author_facet Xu, Dejia
Liang, Hanwen
Bhatt, Neel P.
Hu, Hezhen
Liang, Hanxue
Plataniotis, Konstantinos N.
Wang, Zhangyang
contents Recent advancements in diffusion models for 2D and 3D content creation have sparked a surge of interest in generating 4D content. However, the scarcity of 3D scene datasets constrains current methodologies to primarily object-centric generation. To overcome this limitation, we present Comp4D, a novel framework for Compositional 4D Generation. Unlike conventional methods that generate a singular 4D representation of the entire scene, Comp4D innovatively constructs each 4D object within the scene separately. Utilizing Large Language Models (LLMs), the framework begins by decomposing an input text prompt into distinct entities and maps out their trajectories. It then constructs the compositional 4D scene by accurately positioning these objects along their designated paths. To refine the scene, our method employs a compositional score distillation technique guided by the pre-defined trajectories, utilizing pre-trained diffusion models across text-to-image, text-to-video, and text-to-3D domains. Extensive experiments demonstrate our outstanding 4D content creation capability compared to prior arts, showcasing superior visual quality, motion fidelity, and enhanced object interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16993
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Comp4D: LLM-Guided Compositional 4D Scene Generation
Xu, Dejia
Liang, Hanwen
Bhatt, Neel P.
Hu, Hezhen
Liang, Hanxue
Plataniotis, Konstantinos N.
Wang, Zhangyang
Computer Vision and Pattern Recognition
Recent advancements in diffusion models for 2D and 3D content creation have sparked a surge of interest in generating 4D content. However, the scarcity of 3D scene datasets constrains current methodologies to primarily object-centric generation. To overcome this limitation, we present Comp4D, a novel framework for Compositional 4D Generation. Unlike conventional methods that generate a singular 4D representation of the entire scene, Comp4D innovatively constructs each 4D object within the scene separately. Utilizing Large Language Models (LLMs), the framework begins by decomposing an input text prompt into distinct entities and maps out their trajectories. It then constructs the compositional 4D scene by accurately positioning these objects along their designated paths. To refine the scene, our method employs a compositional score distillation technique guided by the pre-defined trajectories, utilizing pre-trained diffusion models across text-to-image, text-to-video, and text-to-3D domains. Extensive experiments demonstrate our outstanding 4D content creation capability compared to prior arts, showcasing superior visual quality, motion fidelity, and enhanced object interactions.
title Comp4D: LLM-Guided Compositional 4D Scene Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.16993