Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914306590769152 |
|---|---|
| author | Xiong, Chen He, Zhiyuan Chen, Pin-Yu Ko, Ching-Yun Ho, Tsung-Yi |
| author_facet | Xiong, Chen He, Zhiyuan Chen, Pin-Yu Ko, Ching-Yun Ho, Tsung-Yi |
| contents | Activation steering is a practical post-training model alignment technique to enhance the utility of Large Language Models (LLMs). Prior to deploying a model as a service, developers can steer a pre-trained model toward specific behavioral objectives, such as compliance or instruction adherence, without the need for retraining. This process is as simple as adding a steering vector to the model's internal representations. However, this capability unintentionally introduces critical and under-explored safety risks. We identify a phenomenon termed Steering Externalities, where steering vectors derived from entirely benign datasets-such as those enforcing strict compliance or specific output formats like JSON-inadvertently erode safety guardrails. Experiments reveal that these interventions act as a force multiplier, creating new vulnerabilities to jailbreaks and increasing attack success rates to over 80% on standard benchmarks by bypassing the initial safety alignment. Ultimately, our results expose a critical blind spot in deployment: benign activation steering systematically erodes the "safety margin," rendering models more vulnerable to black-box attacks and proving that inference-time utility improvements must be rigorously audited for unintended safety externalities. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_04896 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models Xiong, Chen He, Zhiyuan Chen, Pin-Yu Ko, Ching-Yun Ho, Tsung-Yi Cryptography and Security Artificial Intelligence Activation steering is a practical post-training model alignment technique to enhance the utility of Large Language Models (LLMs). Prior to deploying a model as a service, developers can steer a pre-trained model toward specific behavioral objectives, such as compliance or instruction adherence, without the need for retraining. This process is as simple as adding a steering vector to the model's internal representations. However, this capability unintentionally introduces critical and under-explored safety risks. We identify a phenomenon termed Steering Externalities, where steering vectors derived from entirely benign datasets-such as those enforcing strict compliance or specific output formats like JSON-inadvertently erode safety guardrails. Experiments reveal that these interventions act as a force multiplier, creating new vulnerabilities to jailbreaks and increasing attack success rates to over 80% on standard benchmarks by bypassing the initial safety alignment. Ultimately, our results expose a critical blind spot in deployment: benign activation steering systematically erodes the "safety margin," rendering models more vulnerable to black-box attacks and proving that inference-time utility improvements must be rigorously audited for unintended safety externalities. |
| title | Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models |
| topic | Cryptography and Security Artificial Intelligence |
| url | https://arxiv.org/abs/2602.04896 |