GenXD: Generating Any 3D and 4D Scenes

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhao, Yuyang, Lin, Chung-Ching, Lin, Kevin, Yan, Zhiwen, Li, Linjie, Yang, Zhengyuan, Wang, Jianfeng, Lee, Gim Hee, Wang, Lijuan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910684999057408
author Zhao, Yuyang
Lin, Chung-Ching
Lin, Kevin
Yan, Zhiwen
Li, Linjie
Yang, Zhengyuan
Wang, Jianfeng
Lee, Gim Hee
Wang, Lijuan
author_facet Zhao, Yuyang
Lin, Chung-Ching
Lin, Kevin
Yan, Zhiwen
Li, Linjie
Yang, Zhengyuan
Wang, Jianfeng
Lee, Gim Hee
Wang, Lijuan
contents Recent developments in 2D visual generation have been remarkably successful. However, 3D and 4D generation remain challenging in real-world applications due to the lack of large-scale 4D data and effective model design. In this paper, we propose to jointly investigate general 3D and 4D generation by leveraging camera and object movements commonly observed in daily life. Due to the lack of real-world 4D data in the community, we first propose a data curation pipeline to obtain camera poses and object motion strength from videos. Based on this pipeline, we introduce a large-scale real-world 4D scene dataset: CamVid-30K. By leveraging all the 3D and 4D data, we develop our framework, GenXD, which allows us to produce any 3D or 4D scene. We propose multiview-temporal modules, which disentangle camera and object movements, to seamlessly learn from both 3D and 4D data. Additionally, GenXD employs masked latent conditions to support a variety of conditioning views. GenXD can generate videos that follow the camera trajectory as well as consistent 3D views that can be lifted into 3D representations. We perform extensive evaluations across various real-world and synthetic datasets, demonstrating GenXD's effectiveness and versatility compared to previous methods in 3D and 4D generation.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02319
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GenXD: Generating Any 3D and 4D Scenes
Zhao, Yuyang
Lin, Chung-Ching
Lin, Kevin
Yan, Zhiwen
Li, Linjie
Yang, Zhengyuan
Wang, Jianfeng
Lee, Gim Hee
Wang, Lijuan
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent developments in 2D visual generation have been remarkably successful. However, 3D and 4D generation remain challenging in real-world applications due to the lack of large-scale 4D data and effective model design. In this paper, we propose to jointly investigate general 3D and 4D generation by leveraging camera and object movements commonly observed in daily life. Due to the lack of real-world 4D data in the community, we first propose a data curation pipeline to obtain camera poses and object motion strength from videos. Based on this pipeline, we introduce a large-scale real-world 4D scene dataset: CamVid-30K. By leveraging all the 3D and 4D data, we develop our framework, GenXD, which allows us to produce any 3D or 4D scene. We propose multiview-temporal modules, which disentangle camera and object movements, to seamlessly learn from both 3D and 4D data. Additionally, GenXD employs masked latent conditions to support a variety of conditioning views. GenXD can generate videos that follow the camera trajectory as well as consistent 3D views that can be lifted into 3D representations. We perform extensive evaluations across various real-world and synthetic datasets, demonstrating GenXD's effectiveness and versatility compared to previous methods in 3D and 4D generation.
title GenXD: Generating Any 3D and 4D Scenes
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.02319