DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sun, Wenqiang, Chen, Shuo, Liu, Fangfu, Chen, Zilong, Duan, Yueqi, Zhang, Jun, Wang, Yikai
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915008757104640
author Sun, Wenqiang
Chen, Shuo
Liu, Fangfu
Chen, Zilong
Duan, Yueqi
Zhang, Jun
Wang, Yikai
author_facet Sun, Wenqiang
Chen, Shuo
Liu, Fangfu
Chen, Zilong
Duan, Yueqi
Zhang, Jun
Wang, Yikai
contents In this paper, we introduce \textbf{DimensionX}, a framework designed to generate photorealistic 3D and 4D scenes from just a single image with video diffusion. Our approach begins with the insight that both the spatial structure of a 3D scene and the temporal evolution of a 4D scene can be effectively represented through sequences of video frames. While recent video diffusion models have shown remarkable success in producing vivid visuals, they face limitations in directly recovering 3D/4D scenes due to limited spatial and temporal controllability during generation. To overcome this, we propose ST-Director, which decouples spatial and temporal factors in video diffusion by learning dimension-aware LoRAs from dimension-variant data. This controllable video diffusion approach enables precise manipulation of spatial structure and temporal dynamics, allowing us to reconstruct both 3D and 4D representations from sequential frames with the combination of spatial and temporal dimensions. Additionally, to bridge the gap between generated videos and real-world scenes, we introduce a trajectory-aware mechanism for 3D generation and an identity-preserving denoising strategy for 4D generation. Extensive experiments on various real-world and synthetic datasets demonstrate that DimensionX achieves superior results in controllable video generation, as well as in 3D and 4D scene generation, compared with previous methods.
format Preprint
id arxiv_https___arxiv_org_abs_2411_04928
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion
Sun, Wenqiang
Chen, Shuo
Liu, Fangfu
Chen, Zilong
Duan, Yueqi
Zhang, Jun
Wang, Yikai
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
In this paper, we introduce \textbf{DimensionX}, a framework designed to generate photorealistic 3D and 4D scenes from just a single image with video diffusion. Our approach begins with the insight that both the spatial structure of a 3D scene and the temporal evolution of a 4D scene can be effectively represented through sequences of video frames. While recent video diffusion models have shown remarkable success in producing vivid visuals, they face limitations in directly recovering 3D/4D scenes due to limited spatial and temporal controllability during generation. To overcome this, we propose ST-Director, which decouples spatial and temporal factors in video diffusion by learning dimension-aware LoRAs from dimension-variant data. This controllable video diffusion approach enables precise manipulation of spatial structure and temporal dynamics, allowing us to reconstruct both 3D and 4D representations from sequential frames with the combination of spatial and temporal dimensions. Additionally, to bridge the gap between generated videos and real-world scenes, we introduce a trajectory-aware mechanism for 3D generation and an identity-preserving denoising strategy for 4D generation. Extensive experiments on various real-world and synthetic datasets demonstrate that DimensionX achieves superior results in controllable video generation, as well as in 3D and 4D scene generation, compared with previous methods.
title DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
url https://arxiv.org/abs/2411.04928