CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chang, Gyusam, Roh, Wonseok, Jang, Sujin, Lee, Dongwook, Ji, Daehyun, Oh, Gyeongrok, Park, Jinsun, Kim, Jinkyu, Kim, Sangpil
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929266928648192
author Chang, Gyusam
Roh, Wonseok
Jang, Sujin
Lee, Dongwook
Ji, Daehyun
Oh, Gyeongrok
Park, Jinsun
Kim, Jinkyu
Kim, Sangpil
author_facet Chang, Gyusam
Roh, Wonseok
Jang, Sujin
Lee, Dongwook
Ji, Daehyun
Oh, Gyeongrok
Park, Jinsun
Kim, Jinkyu
Kim, Sangpil
contents Recent LiDAR-based 3D Object Detection (3DOD) methods show promising results, but they often do not generalize well to target domains outside the source (or training) data distribution. To reduce such domain gaps and thus to make 3DOD models more generalizable, we introduce a novel unsupervised domain adaptation (UDA) method, called CMDA, which (i) leverages visual semantic cues from an image modality (i.e., camera images) as an effective semantic bridge to close the domain gap in the cross-modal Bird's Eye View (BEV) representations. Further, (ii) we also introduce a self-training-based learning strategy, wherein a model is adversarially trained to generate domain-invariant features, which disrupt the discrimination of whether a feature instance comes from a source or an unseen target domain. Overall, our CMDA framework guides the 3DOD model to generate highly informative and domain-adaptive features for novel data distributions. In our extensive experiments with large-scale benchmarks, such as nuScenes, Waymo, and KITTI, those mentioned above provide significant performance gains for UDA tasks, achieving state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03721
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection
Chang, Gyusam
Roh, Wonseok
Jang, Sujin
Lee, Dongwook
Ji, Daehyun
Oh, Gyeongrok
Park, Jinsun
Kim, Jinkyu
Kim, Sangpil
Computer Vision and Pattern Recognition
Recent LiDAR-based 3D Object Detection (3DOD) methods show promising results, but they often do not generalize well to target domains outside the source (or training) data distribution. To reduce such domain gaps and thus to make 3DOD models more generalizable, we introduce a novel unsupervised domain adaptation (UDA) method, called CMDA, which (i) leverages visual semantic cues from an image modality (i.e., camera images) as an effective semantic bridge to close the domain gap in the cross-modal Bird's Eye View (BEV) representations. Further, (ii) we also introduce a self-training-based learning strategy, wherein a model is adversarially trained to generate domain-invariant features, which disrupt the discrimination of whether a feature instance comes from a source or an unseen target domain. Overall, our CMDA framework guides the 3DOD model to generate highly informative and domain-adaptive features for novel data distributions. In our extensive experiments with large-scale benchmarks, such as nuScenes, Waymo, and KITTI, those mentioned above provide significant performance gains for UDA tasks, achieving state-of-the-art performance.
title CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.03721