Depth Anything with Any Prior

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zehan, Chen, Siyu, Yang, Lihe, Wang, Jialei, Zhang, Ziang, Zhao, Hengshuang, Zhao, Zhou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909611379916800
author Wang, Zehan
Chen, Siyu
Yang, Lihe
Wang, Jialei
Zhang, Ziang
Zhao, Hengshuang
Zhao, Zhou
author_facet Wang, Zehan
Chen, Siyu
Yang, Lihe
Wang, Jialei
Zhang, Ziang
Zhao, Hengshuang
Zhao, Zhou
contents This work presents Prior Depth Anything, a framework that combines incomplete but precise metric information in depth measurement with relative but complete geometric structures in depth prediction, generating accurate, dense, and detailed metric depth maps for any scene. To this end, we design a coarse-to-fine pipeline to progressively integrate the two complementary depth sources. First, we introduce pixel-level metric alignment and distance-aware weighting to pre-fill diverse metric priors by explicitly using depth prediction. It effectively narrows the domain gap between prior patterns, enhancing generalization across varying scenarios. Second, we develop a conditioned monocular depth estimation (MDE) model to refine the inherent noise of depth priors. By conditioning on the normalized pre-filled prior and prediction, the model further implicitly merges the two complementary depth sources. Our model showcases impressive zero-shot generalization across depth completion, super-resolution, and inpainting over 7 real-world datasets, matching or even surpassing previous task-specific methods. More importantly, it performs well on challenging, unseen mixed priors and enables test-time improvements by switching prediction models, providing a flexible accuracy-efficiency trade-off while evolving with advancements in MDE models.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10565
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Depth Anything with Any Prior
Wang, Zehan
Chen, Siyu
Yang, Lihe
Wang, Jialei
Zhang, Ziang
Zhao, Hengshuang
Zhao, Zhou
Computer Vision and Pattern Recognition
This work presents Prior Depth Anything, a framework that combines incomplete but precise metric information in depth measurement with relative but complete geometric structures in depth prediction, generating accurate, dense, and detailed metric depth maps for any scene. To this end, we design a coarse-to-fine pipeline to progressively integrate the two complementary depth sources. First, we introduce pixel-level metric alignment and distance-aware weighting to pre-fill diverse metric priors by explicitly using depth prediction. It effectively narrows the domain gap between prior patterns, enhancing generalization across varying scenarios. Second, we develop a conditioned monocular depth estimation (MDE) model to refine the inherent noise of depth priors. By conditioning on the normalized pre-filled prior and prediction, the model further implicitly merges the two complementary depth sources. Our model showcases impressive zero-shot generalization across depth completion, super-resolution, and inpainting over 7 real-world datasets, matching or even surpassing previous task-specific methods. More importantly, it performs well on challenging, unseen mixed priors and enables test-time improvements by switching prediction models, providing a flexible accuracy-efficiency trade-off while evolving with advancements in MDE models.
title Depth Anything with Any Prior
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.10565