Mono3R: Exploiting Monocular Cues for Geometric 3D Reconstruction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Wenyu, Liu, Sidun, Qiao, Peng, Dou, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912333824000000
author Li, Wenyu
Liu, Sidun
Qiao, Peng
Dou, Yong
author_facet Li, Wenyu
Liu, Sidun
Qiao, Peng
Dou, Yong
contents Recent advances in data-driven geometric multi-view 3D reconstruction foundation models (e.g., DUSt3R) have shown remarkable performance across various 3D vision tasks, facilitated by the release of large-scale, high-quality 3D datasets. However, as we observed, constrained by their matching-based principles, the reconstruction quality of existing models suffers significant degradation in challenging regions with limited matching cues, particularly in weakly textured areas and low-light conditions. To mitigate these limitations, we propose to harness the inherent robustness of monocular geometry estimation to compensate for the inherent shortcomings of matching-based methods. Specifically, we introduce a monocular-guided refinement module that integrates monocular geometric priors into multi-view reconstruction frameworks. This integration substantially enhances the robustness of multi-view reconstruction systems, leading to high-quality feed-forward reconstructions. Comprehensive experiments across multiple benchmarks demonstrate that our method achieves substantial improvements in both mutli-view camera pose estimation and point cloud accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13419
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mono3R: Exploiting Monocular Cues for Geometric 3D Reconstruction
Li, Wenyu
Liu, Sidun
Qiao, Peng
Dou, Yong
Computer Vision and Pattern Recognition
Recent advances in data-driven geometric multi-view 3D reconstruction foundation models (e.g., DUSt3R) have shown remarkable performance across various 3D vision tasks, facilitated by the release of large-scale, high-quality 3D datasets. However, as we observed, constrained by their matching-based principles, the reconstruction quality of existing models suffers significant degradation in challenging regions with limited matching cues, particularly in weakly textured areas and low-light conditions. To mitigate these limitations, we propose to harness the inherent robustness of monocular geometry estimation to compensate for the inherent shortcomings of matching-based methods. Specifically, we introduce a monocular-guided refinement module that integrates monocular geometric priors into multi-view reconstruction frameworks. This integration substantially enhances the robustness of multi-view reconstruction systems, leading to high-quality feed-forward reconstructions. Comprehensive experiments across multiple benchmarks demonstrate that our method achieves substantial improvements in both mutli-view camera pose estimation and point cloud accuracy.
title Mono3R: Exploiting Monocular Cues for Geometric 3D Reconstruction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.13419