PureCLIP-Depth: Prompt-Free and Decoder-Free Monocular Depth Estimation within CLIP Embedding Space

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Miya, Ryutaro, Fushinobu, Kazuyoshi, Kawaguchi, Tatsuya
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918393416777728
author Miya, Ryutaro
Fushinobu, Kazuyoshi
Kawaguchi, Tatsuya
author_facet Miya, Ryutaro
Fushinobu, Kazuyoshi
Kawaguchi, Tatsuya
contents We propose PureCLIP-Depth, a completely prompt-free, decoder-free Monocular Depth Estimation (MDE) model that operates entirely within the Contrastive Language-Image Pre-training (CLIP) embedding space. Unlike recent models that rely heavily on geometric features, we explore a novel approach to MDE driven by conceptual information, performing computations directly within the conceptual CLIP space. The core of our method lies in learning a direct mapping from the RGB domain to the depth domain strictly inside this embedding space. Our approach achieves state-of-the-art performance among CLIP embedding-based models on both indoor and outdoor datasets. The code used in this research is available at: https://github.com/ryutaroLF/PureCLIP-Depth
format Preprint
id arxiv_https___arxiv_org_abs_2603_16238
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PureCLIP-Depth: Prompt-Free and Decoder-Free Monocular Depth Estimation within CLIP Embedding Space
Miya, Ryutaro
Fushinobu, Kazuyoshi
Kawaguchi, Tatsuya
Computer Vision and Pattern Recognition
We propose PureCLIP-Depth, a completely prompt-free, decoder-free Monocular Depth Estimation (MDE) model that operates entirely within the Contrastive Language-Image Pre-training (CLIP) embedding space. Unlike recent models that rely heavily on geometric features, we explore a novel approach to MDE driven by conceptual information, performing computations directly within the conceptual CLIP space. The core of our method lies in learning a direct mapping from the RGB domain to the depth domain strictly inside this embedding space. Our approach achieves state-of-the-art performance among CLIP embedding-based models on both indoor and outdoor datasets. The code used in this research is available at: https://github.com/ryutaroLF/PureCLIP-Depth
title PureCLIP-Depth: Prompt-Free and Decoder-Free Monocular Depth Estimation within CLIP Embedding Space
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.16238