Not All Prompts Are Secure: A Switchable Backdoor Attack Against Pre-trained Vision Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Sheng, Bai, Jiawang, Gao, Kuofeng, Yang, Yong, Li, Yiming, Xia, Shu-tao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910450607718400
author Yang, Sheng
Bai, Jiawang
Gao, Kuofeng
Yang, Yong
Li, Yiming
Xia, Shu-tao
author_facet Yang, Sheng
Bai, Jiawang
Gao, Kuofeng
Yang, Yong
Li, Yiming
Xia, Shu-tao
contents Given the power of vision transformers, a new learning paradigm, pre-training and then prompting, makes it more efficient and effective to address downstream visual recognition tasks. In this paper, we identify a novel security threat towards such a paradigm from the perspective of backdoor attacks. Specifically, an extra prompt token, called the switch token in this work, can turn the backdoor mode on, i.e., converting a benign model into a backdoored one. Once under the backdoor mode, a specific trigger can force the model to predict a target class. It poses a severe risk to the users of cloud API, since the malicious behavior can not be activated and detected under the benign mode, thus making the attack very stealthy. To attack a pre-trained model, our proposed attack, named SWARM, learns a trigger and prompt tokens including a switch token. They are optimized with the clean loss which encourages the model always behaves normally even the trigger presents, and the backdoor loss that ensures the backdoor can be activated by the trigger when the switch is on. Besides, we utilize the cross-mode feature distillation to reduce the effect of the switch token on clean samples. The experiments on diverse visual recognition tasks confirm the success of our switchable backdoor attack, i.e., achieving 95%+ attack success rate, and also being hard to be detected and removed. Our code is available at https://github.com/20000yshust/SWARM.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10612
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Not All Prompts Are Secure: A Switchable Backdoor Attack Against Pre-trained Vision Transformers
Yang, Sheng
Bai, Jiawang
Gao, Kuofeng
Yang, Yong
Li, Yiming
Xia, Shu-tao
Computer Vision and Pattern Recognition
Cryptography and Security
Machine Learning
Given the power of vision transformers, a new learning paradigm, pre-training and then prompting, makes it more efficient and effective to address downstream visual recognition tasks. In this paper, we identify a novel security threat towards such a paradigm from the perspective of backdoor attacks. Specifically, an extra prompt token, called the switch token in this work, can turn the backdoor mode on, i.e., converting a benign model into a backdoored one. Once under the backdoor mode, a specific trigger can force the model to predict a target class. It poses a severe risk to the users of cloud API, since the malicious behavior can not be activated and detected under the benign mode, thus making the attack very stealthy. To attack a pre-trained model, our proposed attack, named SWARM, learns a trigger and prompt tokens including a switch token. They are optimized with the clean loss which encourages the model always behaves normally even the trigger presents, and the backdoor loss that ensures the backdoor can be activated by the trigger when the switch is on. Besides, we utilize the cross-mode feature distillation to reduce the effect of the switch token on clean samples. The experiments on diverse visual recognition tasks confirm the success of our switchable backdoor attack, i.e., achieving 95%+ attack success rate, and also being hard to be detected and removed. Our code is available at https://github.com/20000yshust/SWARM.
title Not All Prompts Are Secure: A Switchable Backdoor Attack Against Pre-trained Vision Transformers
topic Computer Vision and Pattern Recognition
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2405.10612