OA-DET3D: Embedding Object Awareness as a General Plug-in for Multi-Camera 3D Object Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chu, Xiaomeng, Deng, Jiajun, Ji, Jianmin, Zhang, Yu, Li, Houqiang, Zhang, Yanyong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918132740784128
author Chu, Xiaomeng
Deng, Jiajun
Ji, Jianmin
Zhang, Yu
Li, Houqiang
Zhang, Yanyong
author_facet Chu, Xiaomeng
Deng, Jiajun
Ji, Jianmin
Zhang, Yu
Li, Houqiang
Zhang, Yanyong
contents The recent advance in multi-camera 3D object detection is featured by bird's-eye view (BEV) representation or object queries. However, the ill-posed transformation from image-plane view to 3D space inevitably causes feature clutter and distortion, making the objects blur into the background. To this end, we explore how to incorporate supplementary cues for differentiating objects in the transformed feature representation. Formally, we introduce OA-DET3D, a general plug-in module that improves 3D object detection by bringing object awareness into a variety of existing 3D object detection pipelines. Specifically, OA-DET3D boosts the representation of objects by leveraging object-centric depth information and foreground pseudo points. First, we use object-level supervision from the properties of each 3D bounding box to guide the network in learning the depth distribution. Next, we select foreground pixels using a 2D object detector and project them into 3D space for pseudo-voxel feature encoding. Finally, the object-aware depth features and pseudo-voxel features are incorporated into the BEV representation or query feature from the baseline model with a deformable attention mechanism. We conduct extensive experiments on the nuScenes dataset and Argoverse 2 dataset to validate the merits of OA-DET3D. Our method achieves consistent improvements over the BEV-based baselines in terms of both average precision and comprehensive detection score.
format Preprint
id arxiv_https___arxiv_org_abs_2301_05711
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle OA-DET3D: Embedding Object Awareness as a General Plug-in for Multi-Camera 3D Object Detection
Chu, Xiaomeng
Deng, Jiajun
Ji, Jianmin
Zhang, Yu
Li, Houqiang
Zhang, Yanyong
Computer Vision and Pattern Recognition
The recent advance in multi-camera 3D object detection is featured by bird's-eye view (BEV) representation or object queries. However, the ill-posed transformation from image-plane view to 3D space inevitably causes feature clutter and distortion, making the objects blur into the background. To this end, we explore how to incorporate supplementary cues for differentiating objects in the transformed feature representation. Formally, we introduce OA-DET3D, a general plug-in module that improves 3D object detection by bringing object awareness into a variety of existing 3D object detection pipelines. Specifically, OA-DET3D boosts the representation of objects by leveraging object-centric depth information and foreground pseudo points. First, we use object-level supervision from the properties of each 3D bounding box to guide the network in learning the depth distribution. Next, we select foreground pixels using a 2D object detector and project them into 3D space for pseudo-voxel feature encoding. Finally, the object-aware depth features and pseudo-voxel features are incorporated into the BEV representation or query feature from the baseline model with a deformable attention mechanism. We conduct extensive experiments on the nuScenes dataset and Argoverse 2 dataset to validate the merits of OA-DET3D. Our method achieves consistent improvements over the BEV-based baselines in terms of both average precision and comprehensive detection score.
title OA-DET3D: Embedding Object Awareness as a General Plug-in for Multi-Camera 3D Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2301.05711