DISCO: Disentangled Communication Steering for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Torop, Max, Masoomi, Aria, Eskandar, Masih, Dy, Jennifer
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912596396867584
author Torop, Max
Masoomi, Aria
Eskandar, Masih
Dy, Jennifer
author_facet Torop, Max
Masoomi, Aria
Eskandar, Masih
Dy, Jennifer
contents A variety of recent methods guide large language model outputs via the inference-time addition of steering vectors to residual-stream or attention-head representations. In contrast, we propose to inject steering vectors directly into the query and value representation spaces within attention heads. We provide evidence that a greater portion of these spaces exhibit high linear discriminability of concepts --a key property motivating the use of steering vectors-- than attention head outputs. We analytically characterize the effect of our method, which we term DISentangled COmmunication (DISCO) Steering, on attention head outputs. Our analysis reveals that DISCO disentangles a strong but underutilized baseline, steering attention inputs, which implicitly modifies queries and values in a rigid manner. In contrast, DISCO's direct modulation of these components enables more granular control. We find that DISCO achieves superior performance over a number of steering vector baselines across multiple datasets on LLaMA 3.1 8B and Gemma 2 9B, with steering efficacy scoring up to 19.1% higher than the runner-up. Our results support the conclusion that the query and value spaces are powerful building blocks for steering vector methods.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16820
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DISCO: Disentangled Communication Steering for Large Language Models
Torop, Max
Masoomi, Aria
Eskandar, Masih
Dy, Jennifer
Machine Learning
A variety of recent methods guide large language model outputs via the inference-time addition of steering vectors to residual-stream or attention-head representations. In contrast, we propose to inject steering vectors directly into the query and value representation spaces within attention heads. We provide evidence that a greater portion of these spaces exhibit high linear discriminability of concepts --a key property motivating the use of steering vectors-- than attention head outputs. We analytically characterize the effect of our method, which we term DISentangled COmmunication (DISCO) Steering, on attention head outputs. Our analysis reveals that DISCO disentangles a strong but underutilized baseline, steering attention inputs, which implicitly modifies queries and values in a rigid manner. In contrast, DISCO's direct modulation of these components enables more granular control. We find that DISCO achieves superior performance over a number of steering vector baselines across multiple datasets on LLaMA 3.1 8B and Gemma 2 9B, with steering efficacy scoring up to 19.1% higher than the runner-up. Our results support the conclusion that the query and value spaces are powerful building blocks for steering vector methods.
title DISCO: Disentangled Communication Steering for Large Language Models
topic Machine Learning
url https://arxiv.org/abs/2509.16820