Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Pengzhou, Wu, Zongru, Du, Wei, Zhao, Haodong, Lu, Wei, Liu, Gongshen
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929613913980928
author Cheng, Pengzhou
Wu, Zongru
Du, Wei
Zhao, Haodong
Lu, Wei
Liu, Gongshen
author_facet Cheng, Pengzhou
Wu, Zongru
Du, Wei
Zhao, Haodong
Lu, Wei
Liu, Gongshen
contents Language Models (LMs) are becoming increasingly popular in real-world applications. Outsourcing model training and data hosting to third-party platforms has become a standard method for reducing costs. In such a situation, the attacker can manipulate the training process or data to inject a backdoor into models. Backdoor attacks are a serious threat where malicious behavior is activated when triggers are present, otherwise, the model operates normally. However, there is still no systematic and comprehensive review of LMs from the attacker's capabilities and purposes on different backdoor attack surfaces. Moreover, there is a shortage of analysis and comparison of the diverse emerging backdoor countermeasures. Therefore, this work aims to provide the NLP community with a timely review of backdoor attacks and countermeasures. According to the attackers' capability and affected stage of the LMs, the attack surfaces are formalized into four categorizations: attacking the pre-trained model with fine-tuning (APMF) or parameter-efficient fine-tuning (APMP), attacking the final model with training (AFMT), and attacking Large Language Models (ALLM). Thus, attacks under each categorization are combed. The countermeasures are categorized into two general classes: sample inspection and model inspection. Thus, we review countermeasures and analyze their advantages and disadvantages. Also, we summarize the benchmark datasets and provide comparable evaluations for representative attacks and defenses. Drawing the insights from the review, we point out the crucial areas for future research on the backdoor, especially soliciting more efficient and practical countermeasures.
format Preprint
id arxiv_https___arxiv_org_abs_2309_06055
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review
Cheng, Pengzhou
Wu, Zongru
Du, Wei
Zhao, Haodong
Lu, Wei
Liu, Gongshen
Cryptography and Security
Language Models (LMs) are becoming increasingly popular in real-world applications. Outsourcing model training and data hosting to third-party platforms has become a standard method for reducing costs. In such a situation, the attacker can manipulate the training process or data to inject a backdoor into models. Backdoor attacks are a serious threat where malicious behavior is activated when triggers are present, otherwise, the model operates normally. However, there is still no systematic and comprehensive review of LMs from the attacker's capabilities and purposes on different backdoor attack surfaces. Moreover, there is a shortage of analysis and comparison of the diverse emerging backdoor countermeasures. Therefore, this work aims to provide the NLP community with a timely review of backdoor attacks and countermeasures. According to the attackers' capability and affected stage of the LMs, the attack surfaces are formalized into four categorizations: attacking the pre-trained model with fine-tuning (APMF) or parameter-efficient fine-tuning (APMP), attacking the final model with training (AFMT), and attacking Large Language Models (ALLM). Thus, attacks under each categorization are combed. The countermeasures are categorized into two general classes: sample inspection and model inspection. Thus, we review countermeasures and analyze their advantages and disadvantages. Also, we summarize the benchmark datasets and provide comparable evaluations for representative attacks and defenses. Drawing the insights from the review, we point out the crucial areas for future research on the backdoor, especially soliciting more efficient and practical countermeasures.
title Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review
topic Cryptography and Security
url https://arxiv.org/abs/2309.06055