Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiong, Chen, He, Zhiyuan, Chen, Pin-Yu, Ko, Ching-Yun, Ho, Tsung-Yi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914306590769152
author Xiong, Chen
He, Zhiyuan
Chen, Pin-Yu
Ko, Ching-Yun
Ho, Tsung-Yi
author_facet Xiong, Chen
He, Zhiyuan
Chen, Pin-Yu
Ko, Ching-Yun
Ho, Tsung-Yi
contents Activation steering is a practical post-training model alignment technique to enhance the utility of Large Language Models (LLMs). Prior to deploying a model as a service, developers can steer a pre-trained model toward specific behavioral objectives, such as compliance or instruction adherence, without the need for retraining. This process is as simple as adding a steering vector to the model's internal representations. However, this capability unintentionally introduces critical and under-explored safety risks. We identify a phenomenon termed Steering Externalities, where steering vectors derived from entirely benign datasets-such as those enforcing strict compliance or specific output formats like JSON-inadvertently erode safety guardrails. Experiments reveal that these interventions act as a force multiplier, creating new vulnerabilities to jailbreaks and increasing attack success rates to over 80% on standard benchmarks by bypassing the initial safety alignment. Ultimately, our results expose a critical blind spot in deployment: benign activation steering systematically erodes the "safety margin," rendering models more vulnerable to black-box attacks and proving that inference-time utility improvements must be rigorously audited for unintended safety externalities.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04896
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models
Xiong, Chen
He, Zhiyuan
Chen, Pin-Yu
Ko, Ching-Yun
Ho, Tsung-Yi
Cryptography and Security
Artificial Intelligence
Activation steering is a practical post-training model alignment technique to enhance the utility of Large Language Models (LLMs). Prior to deploying a model as a service, developers can steer a pre-trained model toward specific behavioral objectives, such as compliance or instruction adherence, without the need for retraining. This process is as simple as adding a steering vector to the model's internal representations. However, this capability unintentionally introduces critical and under-explored safety risks. We identify a phenomenon termed Steering Externalities, where steering vectors derived from entirely benign datasets-such as those enforcing strict compliance or specific output formats like JSON-inadvertently erode safety guardrails. Experiments reveal that these interventions act as a force multiplier, creating new vulnerabilities to jailbreaks and increasing attack success rates to over 80% on standard benchmarks by bypassing the initial safety alignment. Ultimately, our results expose a critical blind spot in deployment: benign activation steering systematically erodes the "safety margin," rendering models more vulnerable to black-box attacks and proving that inference-time utility improvements must be rigorously audited for unintended safety externalities.
title Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2602.04896