Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sinii, Viacheslav, Balagansky, Nikita, Gerasimov, Gleb, Laptev, Daniil, Aksenov, Yaroslav, Kurochkin, Vadim, Gorbatovski, Alexey, Shaposhnikov, Boris, Gavrilov, Daniil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908803778215936
author Sinii, Viacheslav
Balagansky, Nikita
Gerasimov, Gleb
Laptev, Daniil
Aksenov, Yaroslav
Kurochkin, Vadim
Gorbatovski, Alexey
Shaposhnikov, Boris
Gavrilov, Daniil
author_facet Sinii, Viacheslav
Balagansky, Nikita
Gerasimov, Gleb
Laptev, Daniil
Aksenov, Yaroslav
Kurochkin, Vadim
Gorbatovski, Alexey
Shaposhnikov, Boris
Gavrilov, Daniil
contents The mechanisms by which reasoning training reshapes LLMs' internal computations remain unclear. We study lightweight steering vectors inserted into the base model's residual stream and trained with a reinforcement-learning objective. These vectors explain a large portion of full fine-tuning performance increase while preserving the interpretability of small, additive interventions. We find that (i) the last-layer steering vector acts like a token-substitution bias concentrated on the first generated token, consistently boosting tokens such as "To" and "Step"; (ii) the penultimate-layer vector leaves attention patterns largely intact and instead operates through the MLP and unembedding, preferentially up-weighting process words and structure symbols; and (iii) the steering vectors transfer to other models from the same family. Taken together, these results deepen understanding of how trained steering vectors shape computation and should inform future work in activation engineering and the study of reasoning models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06608
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
Sinii, Viacheslav
Balagansky, Nikita
Gerasimov, Gleb
Laptev, Daniil
Aksenov, Yaroslav
Kurochkin, Vadim
Gorbatovski, Alexey
Shaposhnikov, Boris
Gavrilov, Daniil
Machine Learning
The mechanisms by which reasoning training reshapes LLMs' internal computations remain unclear. We study lightweight steering vectors inserted into the base model's residual stream and trained with a reinforcement-learning objective. These vectors explain a large portion of full fine-tuning performance increase while preserving the interpretability of small, additive interventions. We find that (i) the last-layer steering vector acts like a token-substitution bias concentrated on the first generated token, consistently boosting tokens such as "To" and "Step"; (ii) the penultimate-layer vector leaves attention patterns largely intact and instead operates through the MLP and unembedding, preferentially up-weighting process words and structure symbols; and (iii) the steering vectors transfer to other models from the same family. Taken together, these results deepen understanding of how trained steering vectors shape computation and should inform future work in activation engineering and the study of reasoning models.
title Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
topic Machine Learning
url https://arxiv.org/abs/2509.06608