Refusal Direction is Universal Across Safety-Aligned Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xinpeng, Wang, Mingyang, Liu, Yihong, Schütze, Hinrich, Plank, Barbara
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910032149348352
author Wang, Xinpeng
Wang, Mingyang
Liu, Yihong
Schütze, Hinrich
Plank, Barbara
author_facet Wang, Xinpeng
Wang, Mingyang
Liu, Yihong
Schütze, Hinrich
Plank, Barbara
contents Refusal mechanisms in large language models (LLMs) are essential for ensuring safety. Recent research has revealed that refusal behavior can be mediated by a single direction in activation space, enabling targeted interventions to bypass refusals. While this is primarily demonstrated in an English-centric context, appropriate refusal behavior is important for any language, but poorly understood. In this paper, we investigate the refusal behavior in LLMs across 14 languages using PolyRefuse, a multilingual safety dataset created by translating malicious and benign English prompts into these languages. We uncover the surprising cross-lingual universality of the refusal direction: a vector extracted from English can bypass refusals in other languages with near-perfect effectiveness, without any additional fine-tuning. Even more remarkably, refusal directions derived from any safety-aligned language transfer seamlessly to others. We attribute this transferability to the parallelism of refusal vectors across languages in the embedding space and identify the underlying mechanism behind cross-lingual jailbreaks. These findings provide actionable insights for building more robust multilingual safety defenses and pave the way for a deeper mechanistic understanding of cross-lingual vulnerabilities in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17306
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Refusal Direction is Universal Across Safety-Aligned Languages
Wang, Xinpeng
Wang, Mingyang
Liu, Yihong
Schütze, Hinrich
Plank, Barbara
Computation and Language
Refusal mechanisms in large language models (LLMs) are essential for ensuring safety. Recent research has revealed that refusal behavior can be mediated by a single direction in activation space, enabling targeted interventions to bypass refusals. While this is primarily demonstrated in an English-centric context, appropriate refusal behavior is important for any language, but poorly understood. In this paper, we investigate the refusal behavior in LLMs across 14 languages using PolyRefuse, a multilingual safety dataset created by translating malicious and benign English prompts into these languages. We uncover the surprising cross-lingual universality of the refusal direction: a vector extracted from English can bypass refusals in other languages with near-perfect effectiveness, without any additional fine-tuning. Even more remarkably, refusal directions derived from any safety-aligned language transfer seamlessly to others. We attribute this transferability to the parallelism of refusal vectors across languages in the embedding space and identify the underlying mechanism behind cross-lingual jailbreaks. These findings provide actionable insights for building more robust multilingual safety defenses and pave the way for a deeper mechanistic understanding of cross-lingual vulnerabilities in LLMs.
title Refusal Direction is Universal Across Safety-Aligned Languages
topic Computation and Language
url https://arxiv.org/abs/2505.17306