Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bochkovskii, Aleksei, Delaunoy, Amaël, Germain, Hugo, Santos, Marcel, Zhou, Yichao, Richter, Stephan R., Koltun, Vladlen
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:https://arxiv.org/abs/2410.02073
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910914704310272
author Bochkovskii, Aleksei
Delaunoy, Amaël
Germain, Hugo
Santos, Marcel
Zhou, Yichao
Richter, Stephan R.
Koltun, Vladlen
author_facet Bochkovskii, Aleksei
Delaunoy, Amaël
Germain, Hugo
Santos, Marcel
Zhou, Yichao
Richter, Stephan R.
Koltun, Vladlen
contents We present a foundation model for zero-shot metric monocular depth estimation. Our model, Depth Pro, synthesizes high-resolution depth maps with unparalleled sharpness and high-frequency details. The predictions are metric, with absolute scale, without relying on the availability of metadata such as camera intrinsics. And the model is fast, producing a 2.25-megapixel depth map in 0.3 seconds on a standard GPU. These characteristics are enabled by a number of technical contributions, including an efficient multi-scale vision transformer for dense prediction, a training protocol that combines real and synthetic datasets to achieve high metric accuracy alongside fine boundary tracing, dedicated evaluation metrics for boundary accuracy in estimated depth maps, and state-of-the-art focal length estimation from a single image. Extensive experiments analyze specific design choices and demonstrate that Depth Pro outperforms prior work along multiple dimensions. We release code and weights at https://github.com/apple/ml-depth-pro
format Preprint
id arxiv_https___arxiv_org_abs_2410_02073
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Depth Pro: Sharp Monocular Metric Depth in Less Than a Second
Bochkovskii, Aleksei
Delaunoy, Amaël
Germain, Hugo
Santos, Marcel
Zhou, Yichao
Richter, Stephan R.
Koltun, Vladlen
Computer Vision and Pattern Recognition
Machine Learning
We present a foundation model for zero-shot metric monocular depth estimation. Our model, Depth Pro, synthesizes high-resolution depth maps with unparalleled sharpness and high-frequency details. The predictions are metric, with absolute scale, without relying on the availability of metadata such as camera intrinsics. And the model is fast, producing a 2.25-megapixel depth map in 0.3 seconds on a standard GPU. These characteristics are enabled by a number of technical contributions, including an efficient multi-scale vision transformer for dense prediction, a training protocol that combines real and synthetic datasets to achieve high metric accuracy alongside fine boundary tracing, dedicated evaluation metrics for boundary accuracy in estimated depth maps, and state-of-the-art focal length estimation from a single image. Extensive experiments analyze specific design choices and demonstrate that Depth Pro outperforms prior work along multiple dimensions. We release code and weights at https://github.com/apple/ml-depth-pro
title Depth Pro: Sharp Monocular Metric Depth in Less Than a Second
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2410.02073