Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908803778215936 |
|---|---|
| author | Sinii, Viacheslav Balagansky, Nikita Gerasimov, Gleb Laptev, Daniil Aksenov, Yaroslav Kurochkin, Vadim Gorbatovski, Alexey Shaposhnikov, Boris Gavrilov, Daniil |
| author_facet | Sinii, Viacheslav Balagansky, Nikita Gerasimov, Gleb Laptev, Daniil Aksenov, Yaroslav Kurochkin, Vadim Gorbatovski, Alexey Shaposhnikov, Boris Gavrilov, Daniil |
| contents | The mechanisms by which reasoning training reshapes LLMs' internal computations remain unclear. We study lightweight steering vectors inserted into the base model's residual stream and trained with a reinforcement-learning objective. These vectors explain a large portion of full fine-tuning performance increase while preserving the interpretability of small, additive interventions. We find that (i) the last-layer steering vector acts like a token-substitution bias concentrated on the first generated token, consistently boosting tokens such as "To" and "Step"; (ii) the penultimate-layer vector leaves attention patterns largely intact and instead operates through the MLP and unembedding, preferentially up-weighting process words and structure symbols; and (iii) the steering vectors transfer to other models from the same family. Taken together, these results deepen understanding of how trained steering vectors shape computation and should inform future work in activation engineering and the study of reasoning models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_06608 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors Sinii, Viacheslav Balagansky, Nikita Gerasimov, Gleb Laptev, Daniil Aksenov, Yaroslav Kurochkin, Vadim Gorbatovski, Alexey Shaposhnikov, Boris Gavrilov, Daniil Machine Learning The mechanisms by which reasoning training reshapes LLMs' internal computations remain unclear. We study lightweight steering vectors inserted into the base model's residual stream and trained with a reinforcement-learning objective. These vectors explain a large portion of full fine-tuning performance increase while preserving the interpretability of small, additive interventions. We find that (i) the last-layer steering vector acts like a token-substitution bias concentrated on the first generated token, consistently boosting tokens such as "To" and "Step"; (ii) the penultimate-layer vector leaves attention patterns largely intact and instead operates through the MLP and unembedding, preferentially up-weighting process words and structure symbols; and (iii) the steering vectors transfer to other models from the same family. Taken together, these results deepen understanding of how trained steering vectors shape computation and should inform future work in activation engineering and the study of reasoning models. |
| title | Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2509.06608 |