Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shairah, Harethah Abu, Hammoud, Hasan Abed Al Kader, Turkiyyah, George, Ghanem, Bernard
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912557833388032
author Shairah, Harethah Abu
Hammoud, Hasan Abed Al Kader
Turkiyyah, George
Ghanem, Bernard
author_facet Shairah, Harethah Abu
Hammoud, Hasan Abed Al Kader
Turkiyyah, George
Ghanem, Bernard
contents Safety alignment in Large Language Models (LLMs) often involves mediating internal representations to refuse harmful requests. Recent research has demonstrated that these safety mechanisms can be bypassed by ablating or removing specific representational directions within the model. In this paper, we propose the opposite approach: Rank-One Safety Injection (ROSI), a white-box method that amplifies a model's safety alignment by permanently steering its activations toward the refusal-mediating subspace. ROSI operates as a simple, fine-tuning-free rank-one weight modification applied to all residual stream write matrices. The required safety direction can be computed from a small set of harmful and harmless instruction pairs. We show that ROSI consistently increases safety refusal rates - as evaluated by Llama Guard 3 - while preserving the utility of the model on standard benchmarks such as MMLU, HellaSwag, and Arc. Furthermore, we show that ROSI can also re-align 'uncensored' models by amplifying their own latent safety directions, demonstrating its utility as an effective last-mile safety procedure. Our results suggest that targeted, interpretable weight steering is a cheap and potent mechanism to improve LLM safety, complementing more resource-intensive fine-tuning paradigms.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20766
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
Shairah, Harethah Abu
Hammoud, Hasan Abed Al Kader
Turkiyyah, George
Ghanem, Bernard
Computation and Language
Artificial Intelligence
Machine Learning
Safety alignment in Large Language Models (LLMs) often involves mediating internal representations to refuse harmful requests. Recent research has demonstrated that these safety mechanisms can be bypassed by ablating or removing specific representational directions within the model. In this paper, we propose the opposite approach: Rank-One Safety Injection (ROSI), a white-box method that amplifies a model's safety alignment by permanently steering its activations toward the refusal-mediating subspace. ROSI operates as a simple, fine-tuning-free rank-one weight modification applied to all residual stream write matrices. The required safety direction can be computed from a small set of harmful and harmless instruction pairs. We show that ROSI consistently increases safety refusal rates - as evaluated by Llama Guard 3 - while preserving the utility of the model on standard benchmarks such as MMLU, HellaSwag, and Arc. Furthermore, we show that ROSI can also re-align 'uncensored' models by amplifying their own latent safety directions, demonstrating its utility as an effective last-mile safety procedure. Our results suggest that targeted, interpretable weight steering is a cheap and potent mechanism to improve LLM safety, complementing more resource-intensive fine-tuning paradigms.
title Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.20766