Adversarial Evasion Attack Efficiency against Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Vitorino, João, Maia, Eva, Praça, Isabel
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909222307889152
author Vitorino, João
Maia, Eva
Praça, Isabel
author_facet Vitorino, João
Maia, Eva
Praça, Isabel
contents Large Language Models (LLMs) are valuable for text classification, but their vulnerabilities must not be disregarded. They lack robustness against adversarial examples, so it is pertinent to understand the impacts of different types of perturbations, and assess if those attacks could be replicated by common users with a small amount of perturbations and a small number of queries to a deployed LLM. This work presents an analysis of the effectiveness, efficiency, and practicality of three different types of adversarial attacks against five different LLMs in a sentiment classification task. The obtained results demonstrated the very distinct impacts of the word-level and character-level attacks. The word attacks were more effective, but the character and more constrained attacks were more practical and required a reduced number of perturbations and queries. These differences need to be considered during the development of adversarial defense strategies to train more robust LLMs for intelligent text classification applications.
format Preprint
id arxiv_https___arxiv_org_abs_2406_08050
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adversarial Evasion Attack Efficiency against Large Language Models
Vitorino, João
Maia, Eva
Praça, Isabel
Computation and Language
Machine Learning
Large Language Models (LLMs) are valuable for text classification, but their vulnerabilities must not be disregarded. They lack robustness against adversarial examples, so it is pertinent to understand the impacts of different types of perturbations, and assess if those attacks could be replicated by common users with a small amount of perturbations and a small number of queries to a deployed LLM. This work presents an analysis of the effectiveness, efficiency, and practicality of three different types of adversarial attacks against five different LLMs in a sentiment classification task. The obtained results demonstrated the very distinct impacts of the word-level and character-level attacks. The word attacks were more effective, but the character and more constrained attacks were more practical and required a reduced number of perturbations and queries. These differences need to be considered during the development of adversarial defense strategies to train more robust LLMs for intelligent text classification applications.
title Adversarial Evasion Attack Efficiency against Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2406.08050