Steer Like the LLM: Activation Steering that Mimics Prompting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Heyman, Geert, Vandeputte, Frederik
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914531606790144
author Heyman, Geert
Vandeputte, Frederik
author_facet Heyman, Geert
Vandeputte, Frederik
contents Large language models can be steered at inference time through prompting or activation interventions, but activation steering methods often underperform compared to prompt-based approaches. We propose a framework that formulates prompt steering as a form of activation steering and investigates whether distilling successful prompt steering behavior into simpler, interpretable models can close this gap. Our analysis reveals that popular activation steering methods are not faithful to the mechanics of prompt steering, which applies strong interventions on some tokens while barely affecting others. Based on these insights, we introduce Prompt Steering Replacement (PSR) models that estimate token-specific steering coefficients from the activations themselves and are trained to imitate prompt-based interventions. Experiments on three steering benchmarks across multiple language models show that PSR models outperform existing activation steering methods, especially when controlling for high-coherence completions, and also compare favorably to prompting on AxBench and persona steering.
format Preprint
id arxiv_https___arxiv_org_abs_2605_03907
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Steer Like the LLM: Activation Steering that Mimics Prompting
Heyman, Geert
Vandeputte, Frederik
Computation and Language
Artificial Intelligence
Machine Learning
Large language models can be steered at inference time through prompting or activation interventions, but activation steering methods often underperform compared to prompt-based approaches. We propose a framework that formulates prompt steering as a form of activation steering and investigates whether distilling successful prompt steering behavior into simpler, interpretable models can close this gap. Our analysis reveals that popular activation steering methods are not faithful to the mechanics of prompt steering, which applies strong interventions on some tokens while barely affecting others. Based on these insights, we introduce Prompt Steering Replacement (PSR) models that estimate token-specific steering coefficients from the activations themselves and are trained to imitate prompt-based interventions. Experiments on three steering benchmarks across multiple language models show that PSR models outperform existing activation steering methods, especially when controlling for high-coherence completions, and also compare favorably to prompting on AxBench and persona steering.
title Steer Like the LLM: Activation Steering that Mimics Prompting
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.03907