TRIDE: A Text-assisted Radar-Image weather-aware fusion network for Depth Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Huawei, Wang, Zixu, Feng, Hao, Ott, Julius, Servadei, Lorenzo, Wille, Robert
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913995174182912
author Sun, Huawei
Wang, Zixu
Feng, Hao
Ott, Julius
Servadei, Lorenzo
Wille, Robert
author_facet Sun, Huawei
Wang, Zixu
Feng, Hao
Ott, Julius
Servadei, Lorenzo
Wille, Robert
contents Depth estimation, essential for autonomous driving, seeks to interpret the 3D environment surrounding vehicles. The development of radar sensors, known for their cost-efficiency and robustness, has spurred interest in radar-camera fusion-based solutions. However, existing algorithms fuse features from these modalities without accounting for weather conditions, despite radars being known to be more robust than cameras under adverse weather. Additionally, while Vision-Language models have seen rapid advancement, utilizing language descriptions alongside other modalities for depth estimation remains an open challenge. This paper first introduces a text-generation strategy along with feature extraction and fusion techniques that can assist monocular depth estimation pipelines, leading to improved accuracy across different algorithms on the KITTI dataset. Building on this, we propose TRIDE, a radar-camera fusion algorithm that enhances text feature extraction by incorporating radar point information. To address the impact of weather on sensor performance, we introduce a weather-aware fusion block that adaptively adjusts radar weighting based on current weather conditions. Our method, benchmarked on the nuScenes dataset, demonstrates performance gains over the state-of-the-art, achieving a 12.87% improvement in MAE and a 9.08% improvement in RMSE. Code: https://github.com/harborsarah/TRIDE
format Preprint
id arxiv_https___arxiv_org_abs_2508_08038
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TRIDE: A Text-assisted Radar-Image weather-aware fusion network for Depth Estimation
Sun, Huawei
Wang, Zixu
Feng, Hao
Ott, Julius
Servadei, Lorenzo
Wille, Robert
Computer Vision and Pattern Recognition
Depth estimation, essential for autonomous driving, seeks to interpret the 3D environment surrounding vehicles. The development of radar sensors, known for their cost-efficiency and robustness, has spurred interest in radar-camera fusion-based solutions. However, existing algorithms fuse features from these modalities without accounting for weather conditions, despite radars being known to be more robust than cameras under adverse weather. Additionally, while Vision-Language models have seen rapid advancement, utilizing language descriptions alongside other modalities for depth estimation remains an open challenge. This paper first introduces a text-generation strategy along with feature extraction and fusion techniques that can assist monocular depth estimation pipelines, leading to improved accuracy across different algorithms on the KITTI dataset. Building on this, we propose TRIDE, a radar-camera fusion algorithm that enhances text feature extraction by incorporating radar point information. To address the impact of weather on sensor performance, we introduce a weather-aware fusion block that adaptively adjusts radar weighting based on current weather conditions. Our method, benchmarked on the nuScenes dataset, demonstrates performance gains over the state-of-the-art, achieving a 12.87% improvement in MAE and a 9.08% improvement in RMSE. Code: https://github.com/harborsarah/TRIDE
title TRIDE: A Text-assisted Radar-Image weather-aware fusion network for Depth Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.08038