RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chu, Xiaomeng, Deng, Jiajun, You, Guoliang, Duan, Yifan, Li, Houqiang, Zhang, Yanyong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917967250325504
author Chu, Xiaomeng
Deng, Jiajun
You, Guoliang
Duan, Yifan
Li, Houqiang
Zhang, Yanyong
author_facet Chu, Xiaomeng
Deng, Jiajun
You, Guoliang
Duan, Yifan
Li, Houqiang
Zhang, Yanyong
contents We propose Radar-Camera fusion transformer (RaCFormer) to boost the accuracy of 3D object detection by the following insight. The Radar-Camera fusion in outdoor 3D scene perception is capped by the image-to-BEV transformation--if the depth of pixels is not accurately estimated, the naive combination of BEV features actually integrates unaligned visual content. To avoid this problem, we propose a query-based framework that enables adaptive sampling of instance-relevant features from both the bird's-eye view (BEV) and the original image view. Furthermore, we enhance system performance by two key designs: optimizing query initialization and strengthening the representational capacity of BEV. For the former, we introduce an adaptive circular distribution in polar coordinates to refine the initialization of object queries, allowing for a distance-based adjustment of query density. For the latter, we initially incorporate a radar-guided depth head to refine the transformation from image view to BEV. Subsequently, we focus on leveraging the Doppler effect of radar and introduce an implicit dynamic catcher to capture the temporal elements within the BEV. Extensive experiments on nuScenes and View-of-Delft (VoD) datasets validate the merits of our design. Remarkably, our method achieves superior results of 64.9% mAP and 70.2% NDS on nuScenes. RaCFormer also secures the state-of-the-art performance on the VoD dataset. Code is available at https://github.com/cxmomo/RaCFormer.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12725
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera Fusion
Chu, Xiaomeng
Deng, Jiajun
You, Guoliang
Duan, Yifan
Li, Houqiang
Zhang, Yanyong
Computer Vision and Pattern Recognition
We propose Radar-Camera fusion transformer (RaCFormer) to boost the accuracy of 3D object detection by the following insight. The Radar-Camera fusion in outdoor 3D scene perception is capped by the image-to-BEV transformation--if the depth of pixels is not accurately estimated, the naive combination of BEV features actually integrates unaligned visual content. To avoid this problem, we propose a query-based framework that enables adaptive sampling of instance-relevant features from both the bird's-eye view (BEV) and the original image view. Furthermore, we enhance system performance by two key designs: optimizing query initialization and strengthening the representational capacity of BEV. For the former, we introduce an adaptive circular distribution in polar coordinates to refine the initialization of object queries, allowing for a distance-based adjustment of query density. For the latter, we initially incorporate a radar-guided depth head to refine the transformation from image view to BEV. Subsequently, we focus on leveraging the Doppler effect of radar and introduce an implicit dynamic catcher to capture the temporal elements within the BEV. Extensive experiments on nuScenes and View-of-Delft (VoD) datasets validate the merits of our design. Remarkably, our method achieves superior results of 64.9% mAP and 70.2% NDS on nuScenes. RaCFormer also secures the state-of-the-art performance on the VoD dataset. Code is available at https://github.com/cxmomo/RaCFormer.
title RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera Fusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.12725