SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Siu, Vincent, Crispino, Nicholas, Park, David, Henry, Nathan W., Wang, Zhun, Liu, Yang, Song, Dawn, Wang, Chenguang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908597832646656
author Siu, Vincent
Crispino, Nicholas
Park, David
Henry, Nathan W.
Wang, Zhun
Liu, Yang
Song, Dawn
Wang, Chenguang
author_facet Siu, Vincent
Crispino, Nicholas
Park, David
Henry, Nathan W.
Wang, Zhun
Liu, Yang
Song, Dawn
Wang, Chenguang
contents We introduce SteeringSafety, a systematic framework for evaluating representation steering methods across seven safety perspectives spanning 17 datasets. While prior work highlights general capabilities of representation steering, we systematically explore safety perspectives including bias, harmfulness, hallucination, social behaviors, reasoning, epistemic integrity, and normative judgment. Our framework provides modularized building blocks for state-of-the-art steering methods, enabling unified implementation of DIM, ACE, CAA, PCA, and LAT with recent enhancements like conditional steering. Results on Gemma-2-2B, Llama-3.1-8B, and Qwen-2.5-7B reveal that strong steering performance depends critically on pairing of method, model, and specific perspective. DIM shows consistent effectiveness, but all methods exhibit substantial entanglement: social behaviors show highest vulnerability (reaching degradation as high as 76%), jailbreaking often compromises normative judgment, and hallucination steering unpredictably shifts political views. Our findings underscore the critical need for holistic safety evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13450
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
Siu, Vincent
Crispino, Nicholas
Park, David
Henry, Nathan W.
Wang, Zhun
Liu, Yang
Song, Dawn
Wang, Chenguang
Artificial Intelligence
Computation and Language
Machine Learning
We introduce SteeringSafety, a systematic framework for evaluating representation steering methods across seven safety perspectives spanning 17 datasets. While prior work highlights general capabilities of representation steering, we systematically explore safety perspectives including bias, harmfulness, hallucination, social behaviors, reasoning, epistemic integrity, and normative judgment. Our framework provides modularized building blocks for state-of-the-art steering methods, enabling unified implementation of DIM, ACE, CAA, PCA, and LAT with recent enhancements like conditional steering. Results on Gemma-2-2B, Llama-3.1-8B, and Qwen-2.5-7B reveal that strong steering performance depends critically on pairing of method, model, and specific perspective. DIM shows consistent effectiveness, but all methods exhibit substantial entanglement: social behaviors show highest vulnerability (reaching degradation as high as 76%), jailbreaking often compromises normative judgment, and hallucination steering unpredictably shifts political views. Our findings underscore the critical need for holistic safety evaluations.
title SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.13450