Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ferrand, Jean-Charles Noirot, Beugin, Yohan, Pauley, Eric, Sheatsley, Ryan, McDaniel, Patrick
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912910743175168
author Ferrand, Jean-Charles Noirot
Beugin, Yohan
Pauley, Eric
Sheatsley, Ryan
McDaniel, Patrick
author_facet Ferrand, Jean-Charles Noirot
Beugin, Yohan
Pauley, Eric
Sheatsley, Ryan
McDaniel, Patrick
contents Alignment in large language models (LLMs) is used to enforce guidelines such as safety. Yet, alignment fails in the face of jailbreak attacks that modify inputs to induce unsafe outputs. In this paper, we introduce and evaluate a new technique for jailbreak attacks. We observe that alignment embeds a safety classifier in the LLM responsible for deciding between refusal and compliance, and seek to extract an approximation of this classifier: a surrogate classifier. To this end, we build candidate classifiers from subsets of the LLM. We first evaluate the degree to which candidate classifiers approximate the LLM's safety classifier in benign and adversarial settings. Then, we attack the candidates and measure how well the resulting adversarial inputs transfer to the LLM. Our evaluation shows that the best candidates achieve accurate agreement (an F1 score above 80%) using as little as 20% of the model architecture. Further, we find that attacks mounted on the surrogate classifiers can be transferred to the LLM with high success. For example, a surrogate using only 50% of the Llama 2 model achieved an attack success rate (ASR) of 70% with half the memory footprint and runtime -- a substantial improvement over attacking the LLM directly, where we only observed a 22% ASR. These results show that extracting surrogate classifiers is an effective and efficient means for modeling (and therein addressing) the vulnerability of aligned models to jailbreaking attacks. The code is available at https://github.com/jcnf0/targeting-alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2501_16534
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs
Ferrand, Jean-Charles Noirot
Beugin, Yohan
Pauley, Eric
Sheatsley, Ryan
McDaniel, Patrick
Cryptography and Security
Artificial Intelligence
Alignment in large language models (LLMs) is used to enforce guidelines such as safety. Yet, alignment fails in the face of jailbreak attacks that modify inputs to induce unsafe outputs. In this paper, we introduce and evaluate a new technique for jailbreak attacks. We observe that alignment embeds a safety classifier in the LLM responsible for deciding between refusal and compliance, and seek to extract an approximation of this classifier: a surrogate classifier. To this end, we build candidate classifiers from subsets of the LLM. We first evaluate the degree to which candidate classifiers approximate the LLM's safety classifier in benign and adversarial settings. Then, we attack the candidates and measure how well the resulting adversarial inputs transfer to the LLM. Our evaluation shows that the best candidates achieve accurate agreement (an F1 score above 80%) using as little as 20% of the model architecture. Further, we find that attacks mounted on the surrogate classifiers can be transferred to the LLM with high success. For example, a surrogate using only 50% of the Llama 2 model achieved an attack success rate (ASR) of 70% with half the memory footprint and runtime -- a substantial improvement over attacking the LLM directly, where we only observed a 22% ASR. These results show that extracting surrogate classifiers is an effective and efficient means for modeling (and therein addressing) the vulnerability of aligned models to jailbreaking attacks. The code is available at https://github.com/jcnf0/targeting-alignment.
title Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2501.16534