An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Boreiko, Valentyn, Panfilov, Alexander, Voracek, Vaclav, Hein, Matthias, Geiping, Jonas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918054051446784
author Boreiko, Valentyn
Panfilov, Alexander
Voracek, Vaclav
Hein, Matthias
Geiping, Jonas
author_facet Boreiko, Valentyn
Panfilov, Alexander
Voracek, Vaclav
Hein, Matthias
Geiping, Jonas
contents A plethora of jailbreaking attacks have been proposed to obtain harmful responses from safety-tuned LLMs. These methods largely succeed in coercing the target output in their original settings, but their attacks vary substantially in fluency and computational effort. In this work, we propose a unified threat model for the principled comparison of these methods. Our threat model checks if a given jailbreak is likely to occur in the distribution of text. For this, we build an N-gram language model on 1T tokens, which, unlike model-based perplexity, allows for an LLM-agnostic, nonparametric, and inherently interpretable evaluation. We adapt popular attacks to this threat model, and, for the first time, benchmark these attacks on equal footing with it. After an extensive comparison, we find attack success rates against safety-tuned modern models to be lower than previously presented and that attacks based on discrete optimization significantly outperform recent LLM-based attacks. Being inherently interpretable, our threat model allows for a comprehensive analysis and comparison of jailbreak attacks. We find that effective attacks exploit and abuse infrequent bigrams, either selecting the ones absent from real-world text or rare ones, e.g., specific to Reddit or code datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16222
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
Boreiko, Valentyn
Panfilov, Alexander
Voracek, Vaclav
Hein, Matthias
Geiping, Jonas
Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
A plethora of jailbreaking attacks have been proposed to obtain harmful responses from safety-tuned LLMs. These methods largely succeed in coercing the target output in their original settings, but their attacks vary substantially in fluency and computational effort. In this work, we propose a unified threat model for the principled comparison of these methods. Our threat model checks if a given jailbreak is likely to occur in the distribution of text. For this, we build an N-gram language model on 1T tokens, which, unlike model-based perplexity, allows for an LLM-agnostic, nonparametric, and inherently interpretable evaluation. We adapt popular attacks to this threat model, and, for the first time, benchmark these attacks on equal footing with it. After an extensive comparison, we find attack success rates against safety-tuned modern models to be lower than previously presented and that attacks based on discrete optimization significantly outperform recent LLM-based attacks. Being inherently interpretable, our threat model allows for a comprehensive analysis and comparison of jailbreak attacks. We find that effective attacks exploit and abuse infrequent bigrams, either selecting the ones absent from real-world text or rare ones, e.g., specific to Reddit or code datasets.
title An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
topic Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2410.16222