Salvato in:
Dettagli Bibliografici
Autori principali: Duan, Ruxiao, Hong, Erin, Zhao, Dongxu, Turner, Eric, Wong, Alex, Zhou, Yunwen
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:https://arxiv.org/abs/2603.28896
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911555869736960
author Duan, Ruxiao
Hong, Erin
Zhao, Dongxu
Turner, Eric
Wong, Alex
Zhou, Yunwen
author_facet Duan, Ruxiao
Hong, Erin
Zhao, Dongxu
Turner, Eric
Wong, Alex
Zhou, Yunwen
contents Feed-forward foundation models for multi-view 3-dimensional (3D) reconstruction have been trained on large-scale datasets of perspective images; when tested on wide field-of-view images, e.g., from a fisheye camera, their performance degrades. Their error arises from changes in spatial positions of pixels due to a non-linear projection model that maps 3D points onto the 2D image plane. While one may surmise that training on fisheye images would resolve this problem, there are far fewer fisheye images with ground truth than perspective images, which limit generalization. To enable inference on imagery exhibiting high radial distortion, we propose Fisheye3R, a novel adaptation framework that extends these multi-view 3D reconstruction foundation models to natively accommodate fisheye inputs without performance regression on perspective images. To address the scarcity of fisheye images and ground truth, we introduce flexible learning schemes that support self-supervised adaptation using only unlabeled perspective images and supervised adaptation without any fisheye training data. Extensive experiments across three foundation models, including VGGT, $π^3$, and MapAnything, demonstrate that our approach consistently improves camera pose, depth, point map, and field-of-view estimation on fisheye images.
format Preprint
id arxiv_https___arxiv_org_abs_2603_28896
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Fisheye3R: Adapting Unified 3D Feed-Forward Foundation Models to Fisheye Lenses
Duan, Ruxiao
Hong, Erin
Zhao, Dongxu
Turner, Eric
Wong, Alex
Zhou, Yunwen
Computer Vision and Pattern Recognition
Feed-forward foundation models for multi-view 3-dimensional (3D) reconstruction have been trained on large-scale datasets of perspective images; when tested on wide field-of-view images, e.g., from a fisheye camera, their performance degrades. Their error arises from changes in spatial positions of pixels due to a non-linear projection model that maps 3D points onto the 2D image plane. While one may surmise that training on fisheye images would resolve this problem, there are far fewer fisheye images with ground truth than perspective images, which limit generalization. To enable inference on imagery exhibiting high radial distortion, we propose Fisheye3R, a novel adaptation framework that extends these multi-view 3D reconstruction foundation models to natively accommodate fisheye inputs without performance regression on perspective images. To address the scarcity of fisheye images and ground truth, we introduce flexible learning schemes that support self-supervised adaptation using only unlabeled perspective images and supervised adaptation without any fisheye training data. Extensive experiments across three foundation models, including VGGT, $π^3$, and MapAnything, demonstrate that our approach consistently improves camera pose, depth, point map, and field-of-view estimation on fisheye images.
title Fisheye3R: Adapting Unified 3D Feed-Forward Foundation Models to Fisheye Lenses
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.28896