VADet: Multi-frame LiDAR 3D Object Detection using Variable Aggregation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Huang, Chengjie, Abdelzad, Vahdat, Sedwards, Sean, Czarnecki, Krzysztof
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913581861175296
author Huang, Chengjie
Abdelzad, Vahdat
Sedwards, Sean
Czarnecki, Krzysztof
author_facet Huang, Chengjie
Abdelzad, Vahdat
Sedwards, Sean
Czarnecki, Krzysztof
contents Input aggregation is a simple technique used by state-of-the-art LiDAR 3D object detectors to improve detection. However, increasing aggregation is known to have diminishing returns and even performance degradation, due to objects responding differently to the number of aggregated frames. To address this limitation, we propose an efficient adaptive method, which we call Variable Aggregation Detection (VADet). Instead of aggregating the entire scene using a fixed number of frames, VADet performs aggregation per object, with the number of frames determined by an object's observed properties, such as speed and point density. VADet thus reduces the inherent trade-offs of fixed aggregation and is not architecture specific. To demonstrate its benefits, we apply VADet to three popular single-stage detectors and achieve state-of-the-art performance on the Waymo dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13186
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VADet: Multi-frame LiDAR 3D Object Detection using Variable Aggregation
Huang, Chengjie
Abdelzad, Vahdat
Sedwards, Sean
Czarnecki, Krzysztof
Computer Vision and Pattern Recognition
Input aggregation is a simple technique used by state-of-the-art LiDAR 3D object detectors to improve detection. However, increasing aggregation is known to have diminishing returns and even performance degradation, due to objects responding differently to the number of aggregated frames. To address this limitation, we propose an efficient adaptive method, which we call Variable Aggregation Detection (VADet). Instead of aggregating the entire scene using a fixed number of frames, VADet performs aggregation per object, with the number of frames determined by an object's observed properties, such as speed and point density. VADet thus reduces the inherent trade-offs of fixed aggregation and is not architecture specific. To demonstrate its benefits, we apply VADet to three popular single-stage detectors and achieve state-of-the-art performance on the Waymo dataset.
title VADet: Multi-frame LiDAR 3D Object Detection using Variable Aggregation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.13186