Towards Efficient 3D Object Detection in Bird's-Eye-View Space for Autonomous Driving: A Convolutional-Only Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yuxin, Han, Qiang, Yu, Mengying, Jiang, Yuxin, Yeo, Chaikiat, Li, Yiheng, Huang, Zihang, Liu, Nini, Chen, Hsuanhan, Wu, Xiaojun
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916426503159808
author Li, Yuxin
Han, Qiang
Yu, Mengying
Jiang, Yuxin
Yeo, Chaikiat
Li, Yiheng
Huang, Zihang
Liu, Nini
Chen, Hsuanhan
Wu, Xiaojun
author_facet Li, Yuxin
Han, Qiang
Yu, Mengying
Jiang, Yuxin
Yeo, Chaikiat
Li, Yiheng
Huang, Zihang
Liu, Nini
Chen, Hsuanhan
Wu, Xiaojun
contents 3D object detection in Bird's-Eye-View (BEV) space has recently emerged as a prevalent approach in the field of autonomous driving. Despite the demonstrated improvements in accuracy and velocity estimation compared to perspective view methods, the deployment of BEV-based techniques in real-world autonomous vehicles remains challenging. This is primarily due to their reliance on vision-transformer (ViT) based architectures, which introduce quadratic complexity with respect to the input resolution. To address this issue, we propose an efficient BEV-based 3D detection framework called BEVENet, which leverages a convolutional-only architectural design to circumvent the limitations of ViT models while maintaining the effectiveness of BEV-based methods. Our experiments show that BEVENet is 3$\times$ faster than contemporary state-of-the-art (SOTA) approaches on the NuScenes challenge, achieving a mean average precision (mAP) of 0.456 and a nuScenes detection score (NDS) of 0.555 on the NuScenes validation dataset, with an inference speed of 47.6 frames per second. To the best of our knowledge, this study stands as the first to achieve such significant efficiency improvements for BEV-based methods, highlighting their enhanced feasibility for real-world autonomous driving applications.
format Preprint
id arxiv_https___arxiv_org_abs_2312_00633
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Towards Efficient 3D Object Detection in Bird's-Eye-View Space for Autonomous Driving: A Convolutional-Only Approach
Li, Yuxin
Han, Qiang
Yu, Mengying
Jiang, Yuxin
Yeo, Chaikiat
Li, Yiheng
Huang, Zihang
Liu, Nini
Chen, Hsuanhan
Wu, Xiaojun
Computer Vision and Pattern Recognition
Artificial Intelligence
3D object detection in Bird's-Eye-View (BEV) space has recently emerged as a prevalent approach in the field of autonomous driving. Despite the demonstrated improvements in accuracy and velocity estimation compared to perspective view methods, the deployment of BEV-based techniques in real-world autonomous vehicles remains challenging. This is primarily due to their reliance on vision-transformer (ViT) based architectures, which introduce quadratic complexity with respect to the input resolution. To address this issue, we propose an efficient BEV-based 3D detection framework called BEVENet, which leverages a convolutional-only architectural design to circumvent the limitations of ViT models while maintaining the effectiveness of BEV-based methods. Our experiments show that BEVENet is 3$\times$ faster than contemporary state-of-the-art (SOTA) approaches on the NuScenes challenge, achieving a mean average precision (mAP) of 0.456 and a nuScenes detection score (NDS) of 0.555 on the NuScenes validation dataset, with an inference speed of 47.6 frames per second. To the best of our knowledge, this study stands as the first to achieve such significant efficiency improvements for BEV-based methods, highlighting their enhanced feasibility for real-world autonomous driving applications.
title Towards Efficient 3D Object Detection in Bird's-Eye-View Space for Autonomous Driving: A Convolutional-Only Approach
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2312.00633