The Rogue Scalpel: Activation Steering Compromises LLM Safety

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Korznikov, Anton, Galichin, Andrey, Dontsov, Alexey, Rogov, Oleg Y., Oseledets, Ivan, Tutubalina, Elena
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908835024732160
author Korznikov, Anton
Galichin, Andrey
Dontsov, Alexey
Rogov, Oleg Y.
Oseledets, Ivan
Tutubalina, Elena
author_facet Korznikov, Anton
Galichin, Andrey
Dontsov, Alexey
Rogov, Oleg Y.
Oseledets, Ivan
Tutubalina, Elena
contents Activation steering is a promising technique for controlling LLM behavior by adding semantically meaningful vectors directly into a model's hidden states during inference. It is often framed as a precise, interpretable, and potentially safer alternative to fine-tuning. We demonstrate the opposite: steering systematically breaks model alignment safeguards, making it comply with harmful requests. Through extensive experiments on different model families, we show that even steering in a random direction can increase the probability of harmful compliance from 0% to 1-13%. Alarmingly, steering benign features from a sparse autoencoder (SAE), a common source of interpretable directions, demonstrates a comparable harmful potential. Finally, we show that combining 20 randomly sampled vectors that jailbreak a single prompt creates a universal attack, significantly increasing harmful compliance on unseen requests. These results challenge the paradigm of safety through interpretability, showing that precise control over model internals does not guarantee precise control over model behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22067
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Rogue Scalpel: Activation Steering Compromises LLM Safety
Korznikov, Anton
Galichin, Andrey
Dontsov, Alexey
Rogov, Oleg Y.
Oseledets, Ivan
Tutubalina, Elena
Machine Learning
Artificial Intelligence
Activation steering is a promising technique for controlling LLM behavior by adding semantically meaningful vectors directly into a model's hidden states during inference. It is often framed as a precise, interpretable, and potentially safer alternative to fine-tuning. We demonstrate the opposite: steering systematically breaks model alignment safeguards, making it comply with harmful requests. Through extensive experiments on different model families, we show that even steering in a random direction can increase the probability of harmful compliance from 0% to 1-13%. Alarmingly, steering benign features from a sparse autoencoder (SAE), a common source of interpretable directions, demonstrates a comparable harmful potential. Finally, we show that combining 20 randomly sampled vectors that jailbreak a single prompt creates a universal attack, significantly increasing harmful compliance on unseen requests. These results challenge the paradigm of safety through interpretability, showing that precise control over model internals does not guarantee precise control over model behavior.
title The Rogue Scalpel: Activation Steering Compromises LLM Safety
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.22067