Extending Foundational Monocular Depth Estimators to Fisheye Cameras with Calibration Tokens

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gangopadhyay, Suchisrit, Kim, Jung-Hee, Chen, Xien, Rim, Patrick, Park, Hyoungseob, Wong, Alex
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910053336875008
author Gangopadhyay, Suchisrit
Kim, Jung-Hee
Chen, Xien
Rim, Patrick
Park, Hyoungseob
Wong, Alex
author_facet Gangopadhyay, Suchisrit
Kim, Jung-Hee
Chen, Xien
Rim, Patrick
Park, Hyoungseob
Wong, Alex
contents We propose a method to extend foundational monocular depth estimators (FMDEs), trained on perspective images, to fisheye images. Despite being trained on tens of millions of images, FMDEs are susceptible to the covariate shift introduced by changes in camera calibration (intrinsic, distortion) parameters, leading to erroneous depth estimates. Our method aligns the distribution of latent embeddings encoding fisheye images to those of perspective images, enabling the reuse of FMDEs for fisheye cameras without retraining or finetuning. To this end, we introduce a set of Calibration Tokens as a light-weight adaptation mechanism that modulates the latent embeddings for alignment. By exploiting the already expressive latent space of FMDEs, we posit that modulating their embeddings avoids the negative impact of artifacts and loss introduced in conventional recalibration or map projection to a canonical reference frame in the image space. Our method is self-supervised and does not require fisheye images but leverages publicly available large-scale perspective image datasets. This is done by recalibrating perspective images to fisheye images, and enforcing consistency between their estimates during training. We evaluate our approach with several FMDEs, on both indoors and outdoors, where we consistently improve over state-of-the-art methods using a single set of tokens for both. Code available at: https://github.com/JungHeeKim29/calibration-token.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04928
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Extending Foundational Monocular Depth Estimators to Fisheye Cameras with Calibration Tokens
Gangopadhyay, Suchisrit
Kim, Jung-Hee
Chen, Xien
Rim, Patrick
Park, Hyoungseob
Wong, Alex
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
We propose a method to extend foundational monocular depth estimators (FMDEs), trained on perspective images, to fisheye images. Despite being trained on tens of millions of images, FMDEs are susceptible to the covariate shift introduced by changes in camera calibration (intrinsic, distortion) parameters, leading to erroneous depth estimates. Our method aligns the distribution of latent embeddings encoding fisheye images to those of perspective images, enabling the reuse of FMDEs for fisheye cameras without retraining or finetuning. To this end, we introduce a set of Calibration Tokens as a light-weight adaptation mechanism that modulates the latent embeddings for alignment. By exploiting the already expressive latent space of FMDEs, we posit that modulating their embeddings avoids the negative impact of artifacts and loss introduced in conventional recalibration or map projection to a canonical reference frame in the image space. Our method is self-supervised and does not require fisheye images but leverages publicly available large-scale perspective image datasets. This is done by recalibrating perspective images to fisheye images, and enforcing consistency between their estimates during training. We evaluate our approach with several FMDEs, on both indoors and outdoors, where we consistently improve over state-of-the-art methods using a single set of tokens for both. Code available at: https://github.com/JungHeeKim29/calibration-token.
title Extending Foundational Monocular Depth Estimators to Fisheye Cameras with Calibration Tokens
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.04928