SimpleBEV: Improved LiDAR-Camera Fusion Architecture for 3D Object Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Yun, Gong, Zhan, Zheng, Peiru, Zhu, Hong, Wu, Shaohua
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910689583431680
author Zhao, Yun
Gong, Zhan
Zheng, Peiru
Zhu, Hong
Wu, Shaohua
author_facet Zhao, Yun
Gong, Zhan
Zheng, Peiru
Zhu, Hong
Wu, Shaohua
contents More and more research works fuse the LiDAR and camera information to improve the 3D object detection of the autonomous driving system. Recently, a simple yet effective fusion framework has achieved an excellent detection performance, fusing the LiDAR and camera features in a unified bird's-eye-view (BEV) space. In this paper, we propose a LiDAR-camera fusion framework, named SimpleBEV, for accurate 3D object detection, which follows the BEV-based fusion framework and improves the camera and LiDAR encoders, respectively. Specifically, we perform the camera-based depth estimation using a cascade network and rectify the depth results with the depth information derived from the LiDAR points. Meanwhile, an auxiliary branch that implements the 3D object detection using only the camera-BEV features is introduced to exploit the camera information during the training phase. Besides, we improve the LiDAR feature extractor by fusing the multi-scaled sparse convolutional features. Experimental results demonstrate the effectiveness of our proposed method. Our method achieves 77.6\% NDS accuracy on the nuScenes dataset, showcasing superior performance in the 3D object detection track.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05292
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SimpleBEV: Improved LiDAR-Camera Fusion Architecture for 3D Object Detection
Zhao, Yun
Gong, Zhan
Zheng, Peiru
Zhu, Hong
Wu, Shaohua
Computer Vision and Pattern Recognition
Artificial Intelligence
More and more research works fuse the LiDAR and camera information to improve the 3D object detection of the autonomous driving system. Recently, a simple yet effective fusion framework has achieved an excellent detection performance, fusing the LiDAR and camera features in a unified bird's-eye-view (BEV) space. In this paper, we propose a LiDAR-camera fusion framework, named SimpleBEV, for accurate 3D object detection, which follows the BEV-based fusion framework and improves the camera and LiDAR encoders, respectively. Specifically, we perform the camera-based depth estimation using a cascade network and rectify the depth results with the depth information derived from the LiDAR points. Meanwhile, an auxiliary branch that implements the 3D object detection using only the camera-BEV features is introduced to exploit the camera information during the training phase. Besides, we improve the LiDAR feature extractor by fusing the multi-scaled sparse convolutional features. Experimental results demonstrate the effectiveness of our proposed method. Our method achieves 77.6\% NDS accuracy on the nuScenes dataset, showcasing superior performance in the 3D object detection track.
title SimpleBEV: Improved LiDAR-Camera Fusion Architecture for 3D Object Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.05292