Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Shuai, Wen, Jinming, Tuan, Luu Anh, Zhao, Junbo, Fu, Jie
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909090453651456
author Zhao, Shuai
Wen, Jinming
Tuan, Luu Anh
Zhao, Junbo
Fu, Jie
author_facet Zhao, Shuai
Wen, Jinming
Tuan, Luu Anh
Zhao, Junbo
Fu, Jie
contents The prompt-based learning paradigm, which bridges the gap between pre-training and fine-tuning, achieves state-of-the-art performance on several NLP tasks, particularly in few-shot settings. Despite being widely applied, prompt-based learning is vulnerable to backdoor attacks. Textual backdoor attacks are designed to introduce targeted vulnerabilities into models by poisoning a subset of training samples through trigger injection and label modification. However, they suffer from flaws such as abnormal natural language expressions resulting from the trigger and incorrect labeling of poisoned samples. In this study, we propose ProAttack, a novel and efficient method for performing clean-label backdoor attacks based on the prompt, which uses the prompt itself as a trigger. Our method does not require external triggers and ensures correct labeling of poisoned samples, improving the stealthy nature of the backdoor attack. With extensive experiments on rich-resource and few-shot text classification tasks, we empirically validate ProAttack's competitive performance in textual backdoor attacks. Notably, in the rich-resource setting, ProAttack achieves state-of-the-art attack success rates in the clean-label backdoor attack benchmark without external triggers.
format Preprint
id arxiv_https___arxiv_org_abs_2305_01219
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
Zhao, Shuai
Wen, Jinming
Tuan, Luu Anh
Zhao, Junbo
Fu, Jie
Computation and Language
Artificial Intelligence
The prompt-based learning paradigm, which bridges the gap between pre-training and fine-tuning, achieves state-of-the-art performance on several NLP tasks, particularly in few-shot settings. Despite being widely applied, prompt-based learning is vulnerable to backdoor attacks. Textual backdoor attacks are designed to introduce targeted vulnerabilities into models by poisoning a subset of training samples through trigger injection and label modification. However, they suffer from flaws such as abnormal natural language expressions resulting from the trigger and incorrect labeling of poisoned samples. In this study, we propose ProAttack, a novel and efficient method for performing clean-label backdoor attacks based on the prompt, which uses the prompt itself as a trigger. Our method does not require external triggers and ensures correct labeling of poisoned samples, improving the stealthy nature of the backdoor attack. With extensive experiments on rich-resource and few-shot text classification tasks, we empirically validate ProAttack's competitive performance in textual backdoor attacks. Notably, in the rich-resource setting, ProAttack achieves state-of-the-art attack success rates in the clean-label backdoor attack benchmark without external triggers.
title Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2305.01219