Refusal in Language Models Is Mediated by a Single Direction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arditi, Andy, Obeso, Oscar, Syed, Aaquib, Paleka, Daniel, Panickssery, Nina, Gurnee, Wes, Nanda, Neel
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929569074774016
author Arditi, Andy
Obeso, Oscar
Syed, Aaquib
Paleka, Daniel
Panickssery, Nina
Gurnee, Wes
Nanda, Neel
author_facet Arditi, Andy
Obeso, Oscar
Syed, Aaquib
Paleka, Daniel
Panickssery, Nina
Gurnee, Wes
Nanda, Neel
contents Conversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this refusal behavior is widespread across chat models, its underlying mechanisms remain poorly understood. In this work, we show that refusal is mediated by a one-dimensional subspace, across 13 popular open-source chat models up to 72B parameters in size. Specifically, for each model, we find a single direction such that erasing this direction from the model's residual stream activations prevents it from refusing harmful instructions, while adding this direction elicits refusal on even harmless instructions. Leveraging this insight, we propose a novel white-box jailbreak method that surgically disables refusal with minimal effect on other capabilities. Finally, we mechanistically analyze how adversarial suffixes suppress propagation of the refusal-mediating direction. Our findings underscore the brittleness of current safety fine-tuning methods. More broadly, our work showcases how an understanding of model internals can be leveraged to develop practical methods for controlling model behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11717
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Refusal in Language Models Is Mediated by a Single Direction
Arditi, Andy
Obeso, Oscar
Syed, Aaquib
Paleka, Daniel
Panickssery, Nina
Gurnee, Wes
Nanda, Neel
Machine Learning
Artificial Intelligence
Computation and Language
Conversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this refusal behavior is widespread across chat models, its underlying mechanisms remain poorly understood. In this work, we show that refusal is mediated by a one-dimensional subspace, across 13 popular open-source chat models up to 72B parameters in size. Specifically, for each model, we find a single direction such that erasing this direction from the model's residual stream activations prevents it from refusing harmful instructions, while adding this direction elicits refusal on even harmless instructions. Leveraging this insight, we propose a novel white-box jailbreak method that surgically disables refusal with minimal effect on other capabilities. Finally, we mechanistically analyze how adversarial suffixes suppress propagation of the refusal-mediating direction. Our findings underscore the brittleness of current safety fine-tuning methods. More broadly, our work showcases how an understanding of model internals can be leveraged to develop practical methods for controlling model behavior.
title Refusal in Language Models Is Mediated by a Single Direction
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2406.11717