Improving Generalization Ability for 3D Object Detection by Learning Sparsity-invariant Features

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Hsin-Cheng, Lin, Chung-Yi, Hsu, Winston H.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916597386444800
author Lu, Hsin-Cheng
Lin, Chung-Yi
Hsu, Winston H.
author_facet Lu, Hsin-Cheng
Lin, Chung-Yi
Hsu, Winston H.
contents In autonomous driving, 3D object detection is essential for accurately identifying and tracking objects. Despite the continuous development of various technologies for this task, a significant drawback is observed in most of them-they experience substantial performance degradation when detecting objects in unseen domains. In this paper, we propose a method to improve the generalization ability for 3D object detection on a single domain. We primarily focus on generalizing from a single source domain to target domains with distinct sensor configurations and scene distributions. To learn sparsity-invariant features from a single source domain, we selectively subsample the source data to a specific beam, using confidence scores determined by the current detector to identify the density that holds utmost importance for the detector. Subsequently, we employ the teacher-student framework to align the Bird's Eye View (BEV) features for different point clouds densities. We also utilize feature content alignment (FCA) and graph-based embedding relationship alignment (GERA) to instruct the detector to be domain-agnostic. Extensive experiments demonstrate that our method exhibits superior generalization capabilities compared to other baselines. Furthermore, our approach even outperforms certain domain adaptation methods that can access to the target domain data.
format Preprint
id arxiv_https___arxiv_org_abs_2502_02322
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Generalization Ability for 3D Object Detection by Learning Sparsity-invariant Features
Lu, Hsin-Cheng
Lin, Chung-Yi
Hsu, Winston H.
Computer Vision and Pattern Recognition
Robotics
In autonomous driving, 3D object detection is essential for accurately identifying and tracking objects. Despite the continuous development of various technologies for this task, a significant drawback is observed in most of them-they experience substantial performance degradation when detecting objects in unseen domains. In this paper, we propose a method to improve the generalization ability for 3D object detection on a single domain. We primarily focus on generalizing from a single source domain to target domains with distinct sensor configurations and scene distributions. To learn sparsity-invariant features from a single source domain, we selectively subsample the source data to a specific beam, using confidence scores determined by the current detector to identify the density that holds utmost importance for the detector. Subsequently, we employ the teacher-student framework to align the Bird's Eye View (BEV) features for different point clouds densities. We also utilize feature content alignment (FCA) and graph-based embedding relationship alignment (GERA) to instruct the detector to be domain-agnostic. Extensive experiments demonstrate that our method exhibits superior generalization capabilities compared to other baselines. Furthermore, our approach even outperforms certain domain adaptation methods that can access to the target domain data.
title Improving Generalization Ability for 3D Object Detection by Learning Sparsity-invariant Features
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2502.02322