Leveraging Stable Diffusion for Monocular Depth Estimation via Image Semantic Encoding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xia, Jingming, Cao, Guanqun, Ma, Guang, Luo, Yiben, Li, Qinzhao, Oyekan, John
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917910961717248
author Xia, Jingming
Cao, Guanqun
Ma, Guang
Luo, Yiben
Li, Qinzhao
Oyekan, John
author_facet Xia, Jingming
Cao, Guanqun
Ma, Guang
Luo, Yiben
Li, Qinzhao
Oyekan, John
contents Monocular depth estimation involves predicting depth from a single RGB image and plays a crucial role in applications such as autonomous driving, robotic navigation, 3D reconstruction, etc. Recent advancements in learning-based methods have significantly improved depth estimation performance. Generative models, particularly Stable Diffusion, have shown remarkable potential in recovering fine details and reconstructing missing regions through large-scale training on diverse datasets. However, models like CLIP, which rely on textual embeddings, face limitations in complex outdoor environments where rich context information is needed. These limitations reduce their effectiveness in such challenging scenarios. Here, we propose a novel image-based semantic embedding that extracts contextual information directly from visual features, significantly improving depth prediction in complex environments. Evaluated on the KITTI and Waymo datasets, our method achieves performance comparable to state-of-the-art models while addressing the shortcomings of CLIP embeddings in handling outdoor scenes. By leveraging visual semantics directly, our method demonstrates enhanced robustness and adaptability in depth estimation tasks, showcasing its potential for application to other visual perception tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01666
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Stable Diffusion for Monocular Depth Estimation via Image Semantic Encoding
Xia, Jingming
Cao, Guanqun
Ma, Guang
Luo, Yiben
Li, Qinzhao
Oyekan, John
Computer Vision and Pattern Recognition
Machine Learning
Monocular depth estimation involves predicting depth from a single RGB image and plays a crucial role in applications such as autonomous driving, robotic navigation, 3D reconstruction, etc. Recent advancements in learning-based methods have significantly improved depth estimation performance. Generative models, particularly Stable Diffusion, have shown remarkable potential in recovering fine details and reconstructing missing regions through large-scale training on diverse datasets. However, models like CLIP, which rely on textual embeddings, face limitations in complex outdoor environments where rich context information is needed. These limitations reduce their effectiveness in such challenging scenarios. Here, we propose a novel image-based semantic embedding that extracts contextual information directly from visual features, significantly improving depth prediction in complex environments. Evaluated on the KITTI and Waymo datasets, our method achieves performance comparable to state-of-the-art models while addressing the shortcomings of CLIP embeddings in handling outdoor scenes. By leveraging visual semantics directly, our method demonstrates enhanced robustness and adaptability in depth estimation tasks, showcasing its potential for application to other visual perception tasks.
title Leveraging Stable Diffusion for Monocular Depth Estimation via Image Semantic Encoding
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2502.01666