UniDepth: Universal Monocular Metric Depth Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Piccinelli, Luigi, Yang, Yung-Hsu, Sakaridis, Christos, Segu, Mattia, Li, Siyuan, Van Gool, Luc, Yu, Fisher
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917624649089024
author Piccinelli, Luigi
Yang, Yung-Hsu
Sakaridis, Christos
Segu, Mattia
Li, Siyuan
Van Gool, Luc
Yu, Fisher
author_facet Piccinelli, Luigi
Yang, Yung-Hsu
Sakaridis, Christos
Segu, Mattia
Li, Siyuan
Van Gool, Luc
Yu, Fisher
contents Accurate monocular metric depth estimation (MMDE) is crucial to solving downstream tasks in 3D perception and modeling. However, the remarkable accuracy of recent MMDE methods is confined to their training domains. These methods fail to generalize to unseen domains even in the presence of moderate domain gaps, which hinders their practical applicability. We propose a new model, UniDepth, capable of reconstructing metric 3D scenes from solely single images across domains. Departing from the existing MMDE methods, UniDepth directly predicts metric 3D points from the input image at inference time without any additional information, striving for a universal and flexible MMDE solution. In particular, UniDepth implements a self-promptable camera module predicting dense camera representation to condition depth features. Our model exploits a pseudo-spherical output representation, which disentangles camera and depth representations. In addition, we propose a geometric invariance loss that promotes the invariance of camera-prompted depth features. Thorough evaluations on ten datasets in a zero-shot regime consistently demonstrate the superior performance of UniDepth, even when compared with methods directly trained on the testing domains. Code and models are available at: https://github.com/lpiccinelli-eth/unidepth
format Preprint
id arxiv_https___arxiv_org_abs_2403_18913
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UniDepth: Universal Monocular Metric Depth Estimation
Piccinelli, Luigi
Yang, Yung-Hsu
Sakaridis, Christos
Segu, Mattia
Li, Siyuan
Van Gool, Luc
Yu, Fisher
Computer Vision and Pattern Recognition
Accurate monocular metric depth estimation (MMDE) is crucial to solving downstream tasks in 3D perception and modeling. However, the remarkable accuracy of recent MMDE methods is confined to their training domains. These methods fail to generalize to unseen domains even in the presence of moderate domain gaps, which hinders their practical applicability. We propose a new model, UniDepth, capable of reconstructing metric 3D scenes from solely single images across domains. Departing from the existing MMDE methods, UniDepth directly predicts metric 3D points from the input image at inference time without any additional information, striving for a universal and flexible MMDE solution. In particular, UniDepth implements a self-promptable camera module predicting dense camera representation to condition depth features. Our model exploits a pseudo-spherical output representation, which disentangles camera and depth representations. In addition, we propose a geometric invariance loss that promotes the invariance of camera-prompted depth features. Thorough evaluations on ten datasets in a zero-shot regime consistently demonstrate the superior performance of UniDepth, even when compared with methods directly trained on the testing domains. Code and models are available at: https://github.com/lpiccinelli-eth/unidepth
title UniDepth: Universal Monocular Metric Depth Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.18913