PromptMono: Cross Prompting Attention for Self-Supervised Monocular Depth Estimation in Challenging Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Changhao, Zhang, Guanwen, Cheng, Zhengyun, Zhou, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908288017235968
author Wang, Changhao
Zhang, Guanwen
Cheng, Zhengyun
Zhou, Wei
author_facet Wang, Changhao
Zhang, Guanwen
Cheng, Zhengyun
Zhou, Wei
contents Considerable efforts have been made to improve monocular depth estimation under ideal conditions. However, in challenging environments, monocular depth estimation still faces difficulties. In this paper, we introduce visual prompt learning for predicting depth across different environments within a unified model, and present a self-supervised learning framework called PromptMono. It employs a set of learnable parameters as visual prompts to capture domain-specific knowledge. To integrate prompting information into image representations, a novel gated cross prompting attention (GCPA) module is proposed, which enhances the depth estimation in diverse conditions. We evaluate the proposed PromptMono on the Oxford Robotcar dataset and the nuScenes dataset. Experimental results demonstrate the superior performance of the proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13796
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PromptMono: Cross Prompting Attention for Self-Supervised Monocular Depth Estimation in Challenging Environments
Wang, Changhao
Zhang, Guanwen
Cheng, Zhengyun
Zhou, Wei
Computer Vision and Pattern Recognition
Considerable efforts have been made to improve monocular depth estimation under ideal conditions. However, in challenging environments, monocular depth estimation still faces difficulties. In this paper, we introduce visual prompt learning for predicting depth across different environments within a unified model, and present a self-supervised learning framework called PromptMono. It employs a set of learnable parameters as visual prompts to capture domain-specific knowledge. To integrate prompting information into image representations, a novel gated cross prompting attention (GCPA) module is proposed, which enhances the depth estimation in diverse conditions. We evaluate the proposed PromptMono on the Oxford Robotcar dataset and the nuScenes dataset. Experimental results demonstrate the superior performance of the proposed method.
title PromptMono: Cross Prompting Attention for Self-Supervised Monocular Depth Estimation in Challenging Environments
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.13796