MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tang, Zhenggang, Fan, Yuchen, Wang, Dilin, Xu, Hongyu, Ranjan, Rakesh, Schwing, Alexander, Yan, Zhicheng
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929621252964352
author Tang, Zhenggang
Fan, Yuchen
Wang, Dilin
Xu, Hongyu
Ranjan, Rakesh
Schwing, Alexander
Yan, Zhicheng
author_facet Tang, Zhenggang
Fan, Yuchen
Wang, Dilin
Xu, Hongyu
Ranjan, Rakesh
Schwing, Alexander
Yan, Zhicheng
contents Recent sparse multi-view scene reconstruction advances like DUSt3R and MASt3R no longer require camera calibration and camera pose estimation. However, they only process a pair of views at a time to infer pixel-aligned pointmaps. When dealing with more than two views, a combinatorial number of error prone pairwise reconstructions are usually followed by an expensive global optimization, which often fails to rectify the pairwise reconstruction errors. To handle more views, reduce errors, and improve inference time, we propose the fast single-stage feed-forward network MV-DUSt3R. At its core are multi-view decoder blocks which exchange information across any number of views while considering one reference view. To make our method robust to reference view selection, we further propose MV-DUSt3R+, which employs cross-reference-view blocks to fuse information across different reference view choices. To further enable novel view synthesis, we extend both by adding and jointly training Gaussian splatting heads. Experiments on multi-view stereo reconstruction, multi-view pose estimation, and novel view synthesis confirm that our methods improve significantly upon prior art. Code will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2412_06974
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds
Tang, Zhenggang
Fan, Yuchen
Wang, Dilin
Xu, Hongyu
Ranjan, Rakesh
Schwing, Alexander
Yan, Zhicheng
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent sparse multi-view scene reconstruction advances like DUSt3R and MASt3R no longer require camera calibration and camera pose estimation. However, they only process a pair of views at a time to infer pixel-aligned pointmaps. When dealing with more than two views, a combinatorial number of error prone pairwise reconstructions are usually followed by an expensive global optimization, which often fails to rectify the pairwise reconstruction errors. To handle more views, reduce errors, and improve inference time, we propose the fast single-stage feed-forward network MV-DUSt3R. At its core are multi-view decoder blocks which exchange information across any number of views while considering one reference view. To make our method robust to reference view selection, we further propose MV-DUSt3R+, which employs cross-reference-view blocks to fuse information across different reference view choices. To further enable novel view synthesis, we extend both by adding and jointly training Gaussian splatting heads. Experiments on multi-view stereo reconstruction, multi-view pose estimation, and novel view synthesis confirm that our methods improve significantly upon prior art. Code will be released.
title MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.06974