Invisible Backdoor Attack against Self-supervised Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Hanrong, Wang, Zhenting, Li, Boheng, Lin, Fulin, Han, Tingxu, Jin, Mingyu, Zhan, Chenlu, Du, Mengnan, Wang, Hongwei, Ma, Shiqing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912306115379200
author Zhang, Hanrong
Wang, Zhenting
Li, Boheng
Lin, Fulin
Han, Tingxu
Jin, Mingyu
Zhan, Chenlu
Du, Mengnan
Wang, Hongwei
Ma, Shiqing
author_facet Zhang, Hanrong
Wang, Zhenting
Li, Boheng
Lin, Fulin
Han, Tingxu
Jin, Mingyu
Zhan, Chenlu
Du, Mengnan
Wang, Hongwei
Ma, Shiqing
contents Self-supervised learning (SSL) models are vulnerable to backdoor attacks. Existing backdoor attacks that are effective in SSL often involve noticeable triggers, like colored patches or visible noise, which are vulnerable to human inspection. This paper proposes an imperceptible and effective backdoor attack against self-supervised models. We first find that existing imperceptible triggers designed for supervised learning are less effective in compromising self-supervised models. We then identify this ineffectiveness is attributed to the overlap in distributions between the backdoor and augmented samples used in SSL. Building on this insight, we design an attack using optimized triggers disentangled with the augmented transformation in the SSL, while remaining imperceptible to human vision. Experiments on five datasets and six SSL algorithms demonstrate our attack is highly effective and stealthy. It also has strong resistance to existing backdoor defenses. Our code can be found at https://github.com/Zhang-Henry/INACTIVE.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14672
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Invisible Backdoor Attack against Self-supervised Learning
Zhang, Hanrong
Wang, Zhenting
Li, Boheng
Lin, Fulin
Han, Tingxu
Jin, Mingyu
Zhan, Chenlu
Du, Mengnan
Wang, Hongwei
Ma, Shiqing
Computer Vision and Pattern Recognition
Self-supervised learning (SSL) models are vulnerable to backdoor attacks. Existing backdoor attacks that are effective in SSL often involve noticeable triggers, like colored patches or visible noise, which are vulnerable to human inspection. This paper proposes an imperceptible and effective backdoor attack against self-supervised models. We first find that existing imperceptible triggers designed for supervised learning are less effective in compromising self-supervised models. We then identify this ineffectiveness is attributed to the overlap in distributions between the backdoor and augmented samples used in SSL. Building on this insight, we design an attack using optimized triggers disentangled with the augmented transformation in the SSL, while remaining imperceptible to human vision. Experiments on five datasets and six SSL algorithms demonstrate our attack is highly effective and stealthy. It also has strong resistance to existing backdoor defenses. Our code can be found at https://github.com/Zhang-Henry/INACTIVE.
title Invisible Backdoor Attack against Self-supervised Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.14672