RadarCam-Depth: Radar-Camera Fusion for Depth Estimation with Learned Metric Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Han, Ma, Yukai, Gu, Yaqing, Hu, Kewei, Liu, Yong, Zuo, Xingxing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916165252546560
author Li, Han
Ma, Yukai
Gu, Yaqing
Hu, Kewei
Liu, Yong
Zuo, Xingxing
author_facet Li, Han
Ma, Yukai
Gu, Yaqing
Hu, Kewei
Liu, Yong
Zuo, Xingxing
contents We present a novel approach for metric dense depth estimation based on the fusion of a single-view image and a sparse, noisy Radar point cloud. The direct fusion of heterogeneous Radar and image data, or their encodings, tends to yield dense depth maps with significant artifacts, blurred boundaries, and suboptimal accuracy. To circumvent this issue, we learn to augment versatile and robust monocular depth prediction with the dense metric scale induced from sparse and noisy Radar data. We propose a Radar-Camera framework for highly accurate and fine-detailed dense depth estimation with four stages, including monocular depth prediction, global scale alignment of monocular depth with sparse Radar points, quasi-dense scale estimation through learning the association between Radar points and image patches, and local scale refinement of dense depth using a scale map learner. Our proposed method significantly outperforms the state-of-the-art Radar-Camera depth estimation methods by reducing the mean absolute error (MAE) of depth estimation by 25.6% and 40.2% on the challenging nuScenes dataset and our self-collected ZJU-4DRadarCam dataset, respectively. Our code and dataset will be released at \url{https://github.com/MMOCKING/RadarCam-Depth}.
format Preprint
id arxiv_https___arxiv_org_abs_2401_04325
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RadarCam-Depth: Radar-Camera Fusion for Depth Estimation with Learned Metric Scale
Li, Han
Ma, Yukai
Gu, Yaqing
Hu, Kewei
Liu, Yong
Zuo, Xingxing
Computer Vision and Pattern Recognition
We present a novel approach for metric dense depth estimation based on the fusion of a single-view image and a sparse, noisy Radar point cloud. The direct fusion of heterogeneous Radar and image data, or their encodings, tends to yield dense depth maps with significant artifacts, blurred boundaries, and suboptimal accuracy. To circumvent this issue, we learn to augment versatile and robust monocular depth prediction with the dense metric scale induced from sparse and noisy Radar data. We propose a Radar-Camera framework for highly accurate and fine-detailed dense depth estimation with four stages, including monocular depth prediction, global scale alignment of monocular depth with sparse Radar points, quasi-dense scale estimation through learning the association between Radar points and image patches, and local scale refinement of dense depth using a scale map learner. Our proposed method significantly outperforms the state-of-the-art Radar-Camera depth estimation methods by reducing the mean absolute error (MAE) of depth estimation by 25.6% and 40.2% on the challenging nuScenes dataset and our self-collected ZJU-4DRadarCam dataset, respectively. Our code and dataset will be released at \url{https://github.com/MMOCKING/RadarCam-Depth}.
title RadarCam-Depth: Radar-Camera Fusion for Depth Estimation with Learned Metric Scale
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.04325