BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhenyu, Lin, Haotong, Feng, Jiashi, Wonka, Peter, Kang, Bingyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913950438785024
author Li, Zhenyu
Lin, Haotong
Feng, Jiashi
Wonka, Peter
Kang, Bingyi
author_facet Li, Zhenyu
Lin, Haotong
Feng, Jiashi
Wonka, Peter
Kang, Bingyi
contents Depth estimation is a fundamental task in computer vision with diverse applications. Recent advancements in deep learning have led to powerful depth foundation models (DFMs), yet their evaluation remains challenging due to inconsistencies in existing protocols. Traditional benchmarks rely on alignment-based metrics that introduce biases, favor certain depth representations, and complicate fair comparisons. In this work, we propose BenchDepth, a new benchmark that evaluates DFMs through five carefully selected downstream proxy tasks: depth completion, stereo matching, monocular feed-forward 3D scene reconstruction, SLAM, and vision-language spatial understanding. Unlike conventional evaluation protocols, our approach assesses DFMs based on their practical utility in real-world applications, bypassing problematic alignment procedures. We benchmark eight state-of-the-art DFMs and provide an in-depth analysis of key findings and observations. We hope our work sparks further discussion in the community on best practices for depth model evaluation and paves the way for future research and advancements in depth estimation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15321
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?
Li, Zhenyu
Lin, Haotong
Feng, Jiashi
Wonka, Peter
Kang, Bingyi
Computer Vision and Pattern Recognition
Depth estimation is a fundamental task in computer vision with diverse applications. Recent advancements in deep learning have led to powerful depth foundation models (DFMs), yet their evaluation remains challenging due to inconsistencies in existing protocols. Traditional benchmarks rely on alignment-based metrics that introduce biases, favor certain depth representations, and complicate fair comparisons. In this work, we propose BenchDepth, a new benchmark that evaluates DFMs through five carefully selected downstream proxy tasks: depth completion, stereo matching, monocular feed-forward 3D scene reconstruction, SLAM, and vision-language spatial understanding. Unlike conventional evaluation protocols, our approach assesses DFMs based on their practical utility in real-world applications, bypassing problematic alignment procedures. We benchmark eight state-of-the-art DFMs and provide an in-depth analysis of key findings and observations. We hope our work sparks further discussion in the community on best practices for depth model evaluation and paves the way for future research and advancements in depth estimation.
title BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.15321