LimeAttack: Local Explainable Method for Textual Hard-Label Adversarial Attack

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Hai, Yang, Zhaoqing, Shang, Weiwei, Wu, Yuren
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911753049210880
author Zhu, Hai
Yang, Zhaoqing
Shang, Weiwei
Wu, Yuren
author_facet Zhu, Hai
Yang, Zhaoqing
Shang, Weiwei
Wu, Yuren
contents Natural language processing models are vulnerable to adversarial examples. Previous textual adversarial attacks adopt gradients or confidence scores to calculate word importance ranking and generate adversarial examples. However, this information is unavailable in the real world. Therefore, we focus on a more realistic and challenging setting, named hard-label attack, in which the attacker can only query the model and obtain a discrete prediction label. Existing hard-label attack algorithms tend to initialize adversarial examples by random substitution and then utilize complex heuristic algorithms to optimize the adversarial perturbation. These methods require a lot of model queries and the attack success rate is restricted by adversary initialization. In this paper, we propose a novel hard-label attack algorithm named LimeAttack, which leverages a local explainable method to approximate word importance ranking, and then adopts beam search to find the optimal solution. Extensive experiments show that LimeAttack achieves the better attacking performance compared with existing hard-label attack under the same query budget. In addition, we evaluate the effectiveness of LimeAttack on large language models, and results indicate that adversarial examples remain a significant threat to large language models. The adversarial examples crafted by LimeAttack are highly transferable and effectively improve model robustness in adversarial training.
format Preprint
id arxiv_https___arxiv_org_abs_2308_00319
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LimeAttack: Local Explainable Method for Textual Hard-Label Adversarial Attack
Zhu, Hai
Yang, Zhaoqing
Shang, Weiwei
Wu, Yuren
Computation and Language
Natural language processing models are vulnerable to adversarial examples. Previous textual adversarial attacks adopt gradients or confidence scores to calculate word importance ranking and generate adversarial examples. However, this information is unavailable in the real world. Therefore, we focus on a more realistic and challenging setting, named hard-label attack, in which the attacker can only query the model and obtain a discrete prediction label. Existing hard-label attack algorithms tend to initialize adversarial examples by random substitution and then utilize complex heuristic algorithms to optimize the adversarial perturbation. These methods require a lot of model queries and the attack success rate is restricted by adversary initialization. In this paper, we propose a novel hard-label attack algorithm named LimeAttack, which leverages a local explainable method to approximate word importance ranking, and then adopts beam search to find the optimal solution. Extensive experiments show that LimeAttack achieves the better attacking performance compared with existing hard-label attack under the same query budget. In addition, we evaluate the effectiveness of LimeAttack on large language models, and results indicate that adversarial examples remain a significant threat to large language models. The adversarial examples crafted by LimeAttack are highly transferable and effectively improve model robustness in adversarial training.
title LimeAttack: Local Explainable Method for Textual Hard-Label Adversarial Attack
topic Computation and Language
url https://arxiv.org/abs/2308.00319