Fusion is Not Enough: Single Modal Attacks on Fusion Models for 3D Object Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Zhiyuan, Choi, Hongjun, Liang, James, Feng, Shiwei, Tao, Guanhong, Liu, Dongfang, Zuzak, Michael, Zhang, Xiangyu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916144096477184
author Cheng, Zhiyuan
Choi, Hongjun
Liang, James
Feng, Shiwei
Tao, Guanhong
Liu, Dongfang
Zuzak, Michael
Zhang, Xiangyu
author_facet Cheng, Zhiyuan
Choi, Hongjun
Liang, James
Feng, Shiwei
Tao, Guanhong
Liu, Dongfang
Zuzak, Michael
Zhang, Xiangyu
contents Multi-sensor fusion (MSF) is widely used in autonomous vehicles (AVs) for perception, particularly for 3D object detection with camera and LiDAR sensors. The purpose of fusion is to capitalize on the advantages of each modality while minimizing its weaknesses. Advanced deep neural network (DNN)-based fusion techniques have demonstrated the exceptional and industry-leading performance. Due to the redundant information in multiple modalities, MSF is also recognized as a general defence strategy against adversarial attacks. In this paper, we attack fusion models from the camera modality that is considered to be of lesser importance in fusion but is more affordable for attackers. We argue that the weakest link of fusion models depends on their most vulnerable modality, and propose an attack framework that targets advanced camera-LiDAR fusion-based 3D object detection models through camera-only adversarial attacks. Our approach employs a two-stage optimization-based strategy that first thoroughly evaluates vulnerable image areas under adversarial attacks, and then applies dedicated attack strategies for different fusion models to generate deployable patches. The evaluations with six advanced camera-LiDAR fusion models and one camera-only model indicate that our attacks successfully compromise all of them. Our approach can either decrease the mean average precision (mAP) of detection performance from 0.824 to 0.353, or degrade the detection score of a target object from 0.728 to 0.156, demonstrating the efficacy of our proposed attack framework. Code is available.
format Preprint
id arxiv_https___arxiv_org_abs_2304_14614
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Fusion is Not Enough: Single Modal Attacks on Fusion Models for 3D Object Detection
Cheng, Zhiyuan
Choi, Hongjun
Liang, James
Feng, Shiwei
Tao, Guanhong
Liu, Dongfang
Zuzak, Michael
Zhang, Xiangyu
Computer Vision and Pattern Recognition
Cryptography and Security
Multi-sensor fusion (MSF) is widely used in autonomous vehicles (AVs) for perception, particularly for 3D object detection with camera and LiDAR sensors. The purpose of fusion is to capitalize on the advantages of each modality while minimizing its weaknesses. Advanced deep neural network (DNN)-based fusion techniques have demonstrated the exceptional and industry-leading performance. Due to the redundant information in multiple modalities, MSF is also recognized as a general defence strategy against adversarial attacks. In this paper, we attack fusion models from the camera modality that is considered to be of lesser importance in fusion but is more affordable for attackers. We argue that the weakest link of fusion models depends on their most vulnerable modality, and propose an attack framework that targets advanced camera-LiDAR fusion-based 3D object detection models through camera-only adversarial attacks. Our approach employs a two-stage optimization-based strategy that first thoroughly evaluates vulnerable image areas under adversarial attacks, and then applies dedicated attack strategies for different fusion models to generate deployable patches. The evaluations with six advanced camera-LiDAR fusion models and one camera-only model indicate that our attacks successfully compromise all of them. Our approach can either decrease the mean average precision (mAP) of detection performance from 0.824 to 0.353, or degrade the detection score of a target object from 0.728 to 0.156, demonstrating the efficacy of our proposed attack framework. Code is available.
title Fusion is Not Enough: Single Modal Attacks on Fusion Models for 3D Object Detection
topic Computer Vision and Pattern Recognition
Cryptography and Security
url https://arxiv.org/abs/2304.14614