Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Mengru, Xu, Ziwen, Mao, Shengyu, Deng, Shumin, Tu, Zhaopeng, Chen, Huajun, Zhang, Ningyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916773999149056
author Wang, Mengru
Xu, Ziwen
Mao, Shengyu
Deng, Shumin
Tu, Zhaopeng
Chen, Huajun
Zhang, Ningyu
author_facet Wang, Mengru
Xu, Ziwen
Mao, Shengyu
Deng, Shumin
Tu, Zhaopeng
Chen, Huajun
Zhang, Ningyu
contents Precise control over language model generation is vital for ensuring both safety and reliability. Although prompt engineering and steering are commonly used to intervene in model behaviors, the vast number of parameters in models often results in highly intertwined internal representations. This interdependency can limit control precision and sometimes lead to unintended side effects. Recent research has explored the use of sparse autoencoders (SAE) to disentangle knowledge in high-dimensional spaces for steering. However, these applications have been limited to toy tasks owing to the nontrivial issue of locating atomic knowledge components. In this paper, we propose Steering Target Atoms (STA), a novel method that isolates and manipulates disentangled knowledge components to enhance safety. Comprehensive experiments demonstrate the effectiveness of our approach. Further analysis reveals that steering exhibits superior robustness and flexibility, particularly in adversarial scenarios. We also apply the steering strategy to the large reasoning model, confirming its effectiveness in precise reasoning control.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20322
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms
Wang, Mengru
Xu, Ziwen
Mao, Shengyu
Deng, Shumin
Tu, Zhaopeng
Chen, Huajun
Zhang, Ningyu
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
Machine Learning
Precise control over language model generation is vital for ensuring both safety and reliability. Although prompt engineering and steering are commonly used to intervene in model behaviors, the vast number of parameters in models often results in highly intertwined internal representations. This interdependency can limit control precision and sometimes lead to unintended side effects. Recent research has explored the use of sparse autoencoders (SAE) to disentangle knowledge in high-dimensional spaces for steering. However, these applications have been limited to toy tasks owing to the nontrivial issue of locating atomic knowledge components. In this paper, we propose Steering Target Atoms (STA), a novel method that isolates and manipulates disentangled knowledge components to enhance safety. Comprehensive experiments demonstrate the effectiveness of our approach. Further analysis reveals that steering exhibits superior robustness and flexibility, particularly in adversarial scenarios. We also apply the steering strategy to the large reasoning model, confirming its effectiveness in precise reasoning control.
title Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2505.20322