Adversarial Attacks on Large Language Models Using Regularized Relaxation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chacko, Samuel Jacob, Biswas, Sajib, Islam, Chashi Mahiul, Liza, Fatema Tabassum, Liu, Xiuwen
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909364668858368
author Chacko, Samuel Jacob
Biswas, Sajib
Islam, Chashi Mahiul
Liza, Fatema Tabassum
Liu, Xiuwen
author_facet Chacko, Samuel Jacob
Biswas, Sajib
Islam, Chashi Mahiul
Liza, Fatema Tabassum
Liu, Xiuwen
contents As powerful Large Language Models (LLMs) are now widely used for numerous practical applications, their safety is of critical importance. While alignment techniques have significantly improved overall safety, LLMs remain vulnerable to carefully crafted adversarial inputs. Consequently, adversarial attack methods are extensively used to study and understand these vulnerabilities. However, current attack methods face significant limitations. Those relying on optimizing discrete tokens suffer from limited efficiency, while continuous optimization techniques fail to generate valid tokens from the model's vocabulary, rendering them impractical for real-world applications. In this paper, we propose a novel technique for adversarial attacks that overcomes these limitations by leveraging regularized gradients with continuous optimization methods. Our approach is two orders of magnitude faster than the state-of-the-art greedy coordinate gradient-based method, significantly improving the attack success rate on aligned language models. Moreover, it generates valid tokens, addressing a fundamental limitation of existing continuous optimization methods. We demonstrate the effectiveness of our attack on five state-of-the-art LLMs using four datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19160
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adversarial Attacks on Large Language Models Using Regularized Relaxation
Chacko, Samuel Jacob
Biswas, Sajib
Islam, Chashi Mahiul
Liza, Fatema Tabassum
Liu, Xiuwen
Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
I.2.7
As powerful Large Language Models (LLMs) are now widely used for numerous practical applications, their safety is of critical importance. While alignment techniques have significantly improved overall safety, LLMs remain vulnerable to carefully crafted adversarial inputs. Consequently, adversarial attack methods are extensively used to study and understand these vulnerabilities. However, current attack methods face significant limitations. Those relying on optimizing discrete tokens suffer from limited efficiency, while continuous optimization techniques fail to generate valid tokens from the model's vocabulary, rendering them impractical for real-world applications. In this paper, we propose a novel technique for adversarial attacks that overcomes these limitations by leveraging regularized gradients with continuous optimization methods. Our approach is two orders of magnitude faster than the state-of-the-art greedy coordinate gradient-based method, significantly improving the attack success rate on aligned language models. Moreover, it generates valid tokens, addressing a fundamental limitation of existing continuous optimization methods. We demonstrate the effectiveness of our attack on five state-of-the-art LLMs using four datasets.
title Adversarial Attacks on Large Language Models Using Regularized Relaxation
topic Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
I.2.7
url https://arxiv.org/abs/2410.19160