Unveiling the Depths: A Multi-Modal Fusion Framework for Challenging Scenarios

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Jialei, Liu, Xianming, Jiang, Junjun, Jiang, Kui, Li, Rui, Cheng, Kai, Ji, Xiangyang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910335845269504
author Xu, Jialei
Liu, Xianming
Jiang, Junjun
Jiang, Kui
Li, Rui
Cheng, Kai
Ji, Xiangyang
author_facet Xu, Jialei
Liu, Xianming
Jiang, Junjun
Jiang, Kui
Li, Rui
Cheng, Kai
Ji, Xiangyang
contents Monocular depth estimation from RGB images plays a pivotal role in 3D vision. However, its accuracy can deteriorate in challenging environments such as nighttime or adverse weather conditions. While long-wave infrared cameras offer stable imaging in such challenging conditions, they are inherently low-resolution, lacking rich texture and semantics as delivered by the RGB image. Current methods focus solely on a single modality due to the difficulties to identify and integrate faithful depth cues from both sources. To address these issues, this paper presents a novel approach that identifies and integrates dominant cross-modality depth features with a learning-based framework. Concretely, we independently compute the coarse depth maps with separate networks by fully utilizing the individual depth cues from each modality. As the advantageous depth spreads across both modalities, we propose a novel confidence loss steering a confidence predictor network to yield a confidence map specifying latent potential depth areas. With the resulting confidence map, we propose a multi-modal fusion network that fuses the final depth in an end-to-end manner. Harnessing the proposed pipeline, our method demonstrates the ability of robust depth estimation in a variety of difficult scenarios. Experimental results on the challenging MS$^2$ and ViViD++ datasets demonstrate the effectiveness and robustness of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11826
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unveiling the Depths: A Multi-Modal Fusion Framework for Challenging Scenarios
Xu, Jialei
Liu, Xianming
Jiang, Junjun
Jiang, Kui
Li, Rui
Cheng, Kai
Ji, Xiangyang
Computer Vision and Pattern Recognition
Monocular depth estimation from RGB images plays a pivotal role in 3D vision. However, its accuracy can deteriorate in challenging environments such as nighttime or adverse weather conditions. While long-wave infrared cameras offer stable imaging in such challenging conditions, they are inherently low-resolution, lacking rich texture and semantics as delivered by the RGB image. Current methods focus solely on a single modality due to the difficulties to identify and integrate faithful depth cues from both sources. To address these issues, this paper presents a novel approach that identifies and integrates dominant cross-modality depth features with a learning-based framework. Concretely, we independently compute the coarse depth maps with separate networks by fully utilizing the individual depth cues from each modality. As the advantageous depth spreads across both modalities, we propose a novel confidence loss steering a confidence predictor network to yield a confidence map specifying latent potential depth areas. With the resulting confidence map, we propose a multi-modal fusion network that fuses the final depth in an end-to-end manner. Harnessing the proposed pipeline, our method demonstrates the ability of robust depth estimation in a variety of difficult scenarios. Experimental results on the challenging MS$^2$ and ViViD++ datasets demonstrate the effectiveness and robustness of our method.
title Unveiling the Depths: A Multi-Modal Fusion Framework for Challenging Scenarios
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.11826