DepthCues: Evaluating Monocular Depth Perception in Large Vision Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Danier, Duolikun, Aygün, Mehmet, Li, Changjian, Bilen, Hakan, Mac Aodha, Oisin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908273927520256
author Danier, Duolikun
Aygün, Mehmet
Li, Changjian
Bilen, Hakan
Mac Aodha, Oisin
author_facet Danier, Duolikun
Aygün, Mehmet
Li, Changjian
Bilen, Hakan
Mac Aodha, Oisin
contents Large-scale pre-trained vision models are becoming increasingly prevalent, offering expressive and generalizable visual representations that benefit various downstream tasks. Recent studies on the emergent properties of these models have revealed their high-level geometric understanding, in particular in the context of depth perception. However, it remains unclear how depth perception arises in these models without explicit depth supervision provided during pre-training. To investigate this, we examine whether the monocular depth cues, similar to those used by the human visual system, emerge in these models. We introduce a new benchmark, DepthCues, designed to evaluate depth cue understanding, and present findings across 20 diverse and representative pre-trained vision models. Our analysis shows that human-like depth cues emerge in more recent larger models. We also explore enhancing depth perception in large vision models by fine-tuning on DepthCues, and find that even without dense depth supervision, this improves depth estimation. To support further research, our benchmark and evaluation code will be made publicly available for studying depth perception in vision models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17385
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
Danier, Duolikun
Aygün, Mehmet
Li, Changjian
Bilen, Hakan
Mac Aodha, Oisin
Computer Vision and Pattern Recognition
Large-scale pre-trained vision models are becoming increasingly prevalent, offering expressive and generalizable visual representations that benefit various downstream tasks. Recent studies on the emergent properties of these models have revealed their high-level geometric understanding, in particular in the context of depth perception. However, it remains unclear how depth perception arises in these models without explicit depth supervision provided during pre-training. To investigate this, we examine whether the monocular depth cues, similar to those used by the human visual system, emerge in these models. We introduce a new benchmark, DepthCues, designed to evaluate depth cue understanding, and present findings across 20 diverse and representative pre-trained vision models. Our analysis shows that human-like depth cues emerge in more recent larger models. We also explore enhancing depth perception in large vision models by fine-tuning on DepthCues, and find that even without dense depth supervision, this improves depth estimation. To support further research, our benchmark and evaluation code will be made publicly available for studying depth perception in vision models.
title DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.17385