Safety Alignment Should Be Made More Than Just A Few Attention Heads

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Chao, Zhang, Zefeng, Yue, Juewei, Li, Quangang, Zhang, Chuang, Liu, Tingwen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918131335692288
author Huang, Chao
Zhang, Zefeng
Yue, Juewei
Li, Quangang
Zhang, Chuang
Liu, Tingwen
author_facet Huang, Chao
Zhang, Zefeng
Yue, Juewei
Li, Quangang
Zhang, Chuang
Liu, Tingwen
contents Current safety alignment for large language models(LLMs) continues to present vulnerabilities, given that adversarial prompting can effectively bypass their safety measures.Our investigation shows that these safety mechanisms predominantly depend on a limited subset of attention heads: removing or ablating these heads can severely compromise model safety. To identify and evaluate these safety-critical components, we introduce RDSHA, a targeted ablation method that leverages the model's refusal direction to pinpoint attention heads mostly responsible for safety behaviors. Further analysis shows that existing jailbreak attacks exploit this concentration by selectively bypassing or manipulating these critical attention heads. To address this issue, we propose AHD, a novel training strategy designed to promote the distributed encoding of safety-related behaviors across numerous attention heads. Experimental results demonstrate that AHD successfully distributes safety-related capabilities across more attention heads. Moreover, evaluations under several mainstream jailbreak attacks show that models trained with AHD exhibit considerably stronger safety robustness, while maintaining overall functional utility.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19697
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Safety Alignment Should Be Made More Than Just A Few Attention Heads
Huang, Chao
Zhang, Zefeng
Yue, Juewei
Li, Quangang
Zhang, Chuang
Liu, Tingwen
Cryptography and Security
Artificial Intelligence
Computation and Language
Current safety alignment for large language models(LLMs) continues to present vulnerabilities, given that adversarial prompting can effectively bypass their safety measures.Our investigation shows that these safety mechanisms predominantly depend on a limited subset of attention heads: removing or ablating these heads can severely compromise model safety. To identify and evaluate these safety-critical components, we introduce RDSHA, a targeted ablation method that leverages the model's refusal direction to pinpoint attention heads mostly responsible for safety behaviors. Further analysis shows that existing jailbreak attacks exploit this concentration by selectively bypassing or manipulating these critical attention heads. To address this issue, we propose AHD, a novel training strategy designed to promote the distributed encoding of safety-related behaviors across numerous attention heads. Experimental results demonstrate that AHD successfully distributes safety-related capabilities across more attention heads. Moreover, evaluations under several mainstream jailbreak attacks show that models trained with AHD exhibit considerably stronger safety robustness, while maintaining overall functional utility.
title Safety Alignment Should Be Made More Than Just A Few Attention Heads
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.19697