CRKD: Enhanced Camera-Radar Object Detection with Cross-modality Knowledge Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Lingjun, Song, Jingyu, Skinner, Katherine A.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909154087534592
author Zhao, Lingjun
Song, Jingyu
Skinner, Katherine A.
author_facet Zhao, Lingjun
Song, Jingyu
Skinner, Katherine A.
contents In the field of 3D object detection for autonomous driving, LiDAR-Camera (LC) fusion is the top-performing sensor configuration. Still, LiDAR is relatively high cost, which hinders adoption of this technology for consumer automobiles. Alternatively, camera and radar are commonly deployed on vehicles already on the road today, but performance of Camera-Radar (CR) fusion falls behind LC fusion. In this work, we propose Camera-Radar Knowledge Distillation (CRKD) to bridge the performance gap between LC and CR detectors with a novel cross-modality KD framework. We use the Bird's-Eye-View (BEV) representation as the shared feature space to enable effective knowledge distillation. To accommodate the unique cross-modality KD path, we propose four distillation losses to help the student learn crucial features from the teacher model. We present extensive evaluations on the nuScenes dataset to demonstrate the effectiveness of the proposed CRKD framework. The project page for CRKD is https://song-jingyu.github.io/CRKD.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19104
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CRKD: Enhanced Camera-Radar Object Detection with Cross-modality Knowledge Distillation
Zhao, Lingjun
Song, Jingyu
Skinner, Katherine A.
Computer Vision and Pattern Recognition
Robotics
In the field of 3D object detection for autonomous driving, LiDAR-Camera (LC) fusion is the top-performing sensor configuration. Still, LiDAR is relatively high cost, which hinders adoption of this technology for consumer automobiles. Alternatively, camera and radar are commonly deployed on vehicles already on the road today, but performance of Camera-Radar (CR) fusion falls behind LC fusion. In this work, we propose Camera-Radar Knowledge Distillation (CRKD) to bridge the performance gap between LC and CR detectors with a novel cross-modality KD framework. We use the Bird's-Eye-View (BEV) representation as the shared feature space to enable effective knowledge distillation. To accommodate the unique cross-modality KD path, we propose four distillation losses to help the student learn crucial features from the teacher model. We present extensive evaluations on the nuScenes dataset to demonstrate the effectiveness of the proposed CRKD framework. The project page for CRKD is https://song-jingyu.github.io/CRKD.
title CRKD: Enhanced Camera-Radar Object Detection with Cross-modality Knowledge Distillation
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2403.19104