Learning High-resolution Vector Representation from Multi-Camera Images for 3D Object Detection

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Zhili, Xu, Shuangjie, Ye, Maosheng, Qian, Zian, Zou, Xiaoyi, Yeung, Dit-Yan, Chen, Qifeng
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929429682323456
author Chen, Zhili
Xu, Shuangjie
Ye, Maosheng
Qian, Zian
Zou, Xiaoyi
Yeung, Dit-Yan
Chen, Qifeng
author_facet Chen, Zhili
Xu, Shuangjie
Ye, Maosheng
Qian, Zian
Zou, Xiaoyi
Yeung, Dit-Yan
Chen, Qifeng
contents The Bird's-Eye-View (BEV) representation is a critical factor that directly impacts the 3D object detection performance, but the traditional BEV grid representation induces quadratic computational cost as the spatial resolution grows. To address this limitation, we present a new camera-based 3D object detector with high-resolution vector representation: VectorFormer. The presented high-resolution vector representation is combined with the lower-resolution BEV representation to efficiently exploit 3D geometry from multi-camera images at a high resolution through our two novel modules: vector scattering and gathering. To this end, the learned vector representation with richer scene contexts can serve as the decoding query for final predictions. We conduct extensive experiments on the nuScenes dataset and demonstrate state-of-the-art performance in NDS and inference time. Furthermore, we investigate query-BEV-based methods incorporated with our proposed vector representation and observe a consistent performance improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2407_15354
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning High-resolution Vector Representation from Multi-Camera Images for 3D Object Detection
Chen, Zhili
Xu, Shuangjie
Ye, Maosheng
Qian, Zian
Zou, Xiaoyi
Yeung, Dit-Yan
Chen, Qifeng
Computer Vision and Pattern Recognition
Robotics
The Bird's-Eye-View (BEV) representation is a critical factor that directly impacts the 3D object detection performance, but the traditional BEV grid representation induces quadratic computational cost as the spatial resolution grows. To address this limitation, we present a new camera-based 3D object detector with high-resolution vector representation: VectorFormer. The presented high-resolution vector representation is combined with the lower-resolution BEV representation to efficiently exploit 3D geometry from multi-camera images at a high resolution through our two novel modules: vector scattering and gathering. To this end, the learned vector representation with richer scene contexts can serve as the decoding query for final predictions. We conduct extensive experiments on the nuScenes dataset and demonstrate state-of-the-art performance in NDS and inference time. Furthermore, we investigate query-BEV-based methods incorporated with our proposed vector representation and observe a consistent performance improvement.
title Learning High-resolution Vector Representation from Multi-Camera Images for 3D Object Detection
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2407.15354