Guardado en:
Detalles Bibliográficos
Autores principales: Han, Xinyang, Gao, Zelin, Kanazawa, Angjoo, Goel, Shubham, Gandelsman, Yossi
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:https://arxiv.org/abs/2404.03652
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914741240201216
author Han, Xinyang
Gao, Zelin
Kanazawa, Angjoo
Goel, Shubham
Gandelsman, Yossi
author_facet Han, Xinyang
Gao, Zelin
Kanazawa, Angjoo
Goel, Shubham
Gandelsman, Yossi
contents Humans can infer 3D structure from 2D images of an object based on past experience and improve their 3D understanding as they see more images. Inspired by this behavior, we introduce SAP3D, a system for 3D reconstruction and novel view synthesis from an arbitrary number of unposed images. Given a few unposed images of an object, we adapt a pre-trained view-conditioned diffusion model together with the camera poses of the images via test-time fine-tuning. The adapted diffusion model and the obtained camera poses are then utilized as instance-specific priors for 3D reconstruction and novel view synthesis. We show that as the number of input images increases, the performance of our approach improves, bridging the gap between optimization-based prior-less 3D reconstruction methods and single-image-to-3D diffusion-based methods. We demonstrate our system on real images as well as standard synthetic benchmarks. Our ablation studies confirm that this adaption behavior is key for more accurate 3D understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03652
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The More You See in 2D, the More You Perceive in 3D
Han, Xinyang
Gao, Zelin
Kanazawa, Angjoo
Goel, Shubham
Gandelsman, Yossi
Computer Vision and Pattern Recognition
Humans can infer 3D structure from 2D images of an object based on past experience and improve their 3D understanding as they see more images. Inspired by this behavior, we introduce SAP3D, a system for 3D reconstruction and novel view synthesis from an arbitrary number of unposed images. Given a few unposed images of an object, we adapt a pre-trained view-conditioned diffusion model together with the camera poses of the images via test-time fine-tuning. The adapted diffusion model and the obtained camera poses are then utilized as instance-specific priors for 3D reconstruction and novel view synthesis. We show that as the number of input images increases, the performance of our approach improves, bridging the gap between optimization-based prior-less 3D reconstruction methods and single-image-to-3D diffusion-based methods. We demonstrate our system on real images as well as standard synthetic benchmarks. Our ablation studies confirm that this adaption behavior is key for more accurate 3D understanding.
title The More You See in 2D, the More You Perceive in 3D
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.03652