CDPR: Cross-modal Diffusion with Polarization for Reliable Monocular Depth Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Rongjia, Jia, Tong, Wang, Hao, Li, Xiaofang, Yang, Xiao, Zhang, Zinuo, Liu, Cuiwei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917402862682112
author Yu, Rongjia
Jia, Tong
Wang, Hao
Li, Xiaofang
Yang, Xiao
Zhang, Zinuo
Liu, Cuiwei
author_facet Yu, Rongjia
Jia, Tong
Wang, Hao
Li, Xiaofang
Yang, Xiao
Zhang, Zinuo
Liu, Cuiwei
contents Monocular depth estimation is a fundamental yet challenging task in computer vision, especially under complex conditions such as textureless surfaces, transparency, and specular reflections. Recent diffusion-based approaches have significantly advanced performance by reformulating depth prediction as a denoising process in the latent space. However, existing methods rely solely on RGB inputs, which often lack sufficient cues in challenging regions. In this work, we present CDPR - Cross-modal Diffusion with Polarization for Reliable Monocular Depth Estimation - a novel diffusion-based framework that integrates physically grounded polarization priors to enhance estimation robustness. Specifically, we encode both RGB and polarization (AoLP/DoLP) images into a shared latent space via a pre-trained Variational Autoencoder (VAE), and dynamically fuse multi-modal information through a learnable confidence-aware gating mechanism. This fusion module adaptively suppresses noisy signals in polarization inputs while preserving informative cues, particularly around reflective or transparent surfaces, and provides the integrated latent representation for subsequent monocular depth estimation. Beyond depth estimation, we further verify that our framework can be easily generalized to surface normal prediction with minimal modification, showcasing its scalability to general polarization-guided dense prediction tasks. Experiments on both synthetic and real-world datasets validate that CDPR significantly outperforms RGB-only baselines in challenging regions while maintaining competitive performance in standard scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2604_11097
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CDPR: Cross-modal Diffusion with Polarization for Reliable Monocular Depth Estimation
Yu, Rongjia
Jia, Tong
Wang, Hao
Li, Xiaofang
Yang, Xiao
Zhang, Zinuo
Liu, Cuiwei
Computer Vision and Pattern Recognition
Monocular depth estimation is a fundamental yet challenging task in computer vision, especially under complex conditions such as textureless surfaces, transparency, and specular reflections. Recent diffusion-based approaches have significantly advanced performance by reformulating depth prediction as a denoising process in the latent space. However, existing methods rely solely on RGB inputs, which often lack sufficient cues in challenging regions. In this work, we present CDPR - Cross-modal Diffusion with Polarization for Reliable Monocular Depth Estimation - a novel diffusion-based framework that integrates physically grounded polarization priors to enhance estimation robustness. Specifically, we encode both RGB and polarization (AoLP/DoLP) images into a shared latent space via a pre-trained Variational Autoencoder (VAE), and dynamically fuse multi-modal information through a learnable confidence-aware gating mechanism. This fusion module adaptively suppresses noisy signals in polarization inputs while preserving informative cues, particularly around reflective or transparent surfaces, and provides the integrated latent representation for subsequent monocular depth estimation. Beyond depth estimation, we further verify that our framework can be easily generalized to surface normal prediction with minimal modification, showcasing its scalability to general polarization-guided dense prediction tasks. Experiments on both synthetic and real-world datasets validate that CDPR significantly outperforms RGB-only baselines in challenging regions while maintaining competitive performance in standard scenes.
title CDPR: Cross-modal Diffusion with Polarization for Reliable Monocular Depth Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.11097