GraphBEV: Towards Robust BEV Feature Alignment for Multi-Modal 3D Object Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Ziying, Yang, Lei, Xu, Shaoqing, Liu, Lin, Xu, Dongyang, Jia, Caiyan, Jia, Feiyang, Wang, Li
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929756507734016
author Song, Ziying
Yang, Lei
Xu, Shaoqing
Liu, Lin
Xu, Dongyang
Jia, Caiyan
Jia, Feiyang
Wang, Li
author_facet Song, Ziying
Yang, Lei
Xu, Shaoqing
Liu, Lin
Xu, Dongyang
Jia, Caiyan
Jia, Feiyang
Wang, Li
contents Integrating LiDAR and camera information into Bird's-Eye-View (BEV) representation has emerged as a crucial aspect of 3D object detection in autonomous driving. However, existing methods are susceptible to the inaccurate calibration relationship between LiDAR and the camera sensor. Such inaccuracies result in errors in depth estimation for the camera branch, ultimately causing misalignment between LiDAR and camera BEV features. In this work, we propose a robust fusion framework called Graph BEV. Addressing errors caused by inaccurate point cloud projection, we introduce a Local Align module that employs neighbor-aware depth features via Graph matching. Additionally, we propose a Global Align module to rectify the misalignment between LiDAR and camera BEV features. Our Graph BEV framework achieves state-of-the-art performance, with an mAP of 70.1\%, surpassing BEV Fusion by 1.6\% on the nuscenes validation set. Importantly, our Graph BEV outperforms BEV Fusion by 8.3\% under conditions with misalignment noise.
format Preprint
id arxiv_https___arxiv_org_abs_2403_11848
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GraphBEV: Towards Robust BEV Feature Alignment for Multi-Modal 3D Object Detection
Song, Ziying
Yang, Lei
Xu, Shaoqing
Liu, Lin
Xu, Dongyang
Jia, Caiyan
Jia, Feiyang
Wang, Li
Computer Vision and Pattern Recognition
Integrating LiDAR and camera information into Bird's-Eye-View (BEV) representation has emerged as a crucial aspect of 3D object detection in autonomous driving. However, existing methods are susceptible to the inaccurate calibration relationship between LiDAR and the camera sensor. Such inaccuracies result in errors in depth estimation for the camera branch, ultimately causing misalignment between LiDAR and camera BEV features. In this work, we propose a robust fusion framework called Graph BEV. Addressing errors caused by inaccurate point cloud projection, we introduce a Local Align module that employs neighbor-aware depth features via Graph matching. Additionally, we propose a Global Align module to rectify the misalignment between LiDAR and camera BEV features. Our Graph BEV framework achieves state-of-the-art performance, with an mAP of 70.1\%, surpassing BEV Fusion by 1.6\% on the nuscenes validation set. Importantly, our Graph BEV outperforms BEV Fusion by 8.3\% under conditions with misalignment noise.
title GraphBEV: Towards Robust BEV Feature Alignment for Multi-Modal 3D Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.11848