Steering LLM Reasoning Through Bias-Only Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918152001028096 |
|---|---|
| author | Sinii, Viacheslav Gorbatovski, Alexey Cherepanov, Artem Shaposhnikov, Boris Balagansky, Nikita Gavrilov, Daniil |
| author_facet | Sinii, Viacheslav Gorbatovski, Alexey Cherepanov, Artem Shaposhnikov, Boris Balagansky, Nikita Gavrilov, Daniil |
| contents | We show that training a single $d$-dimensional steering vector per layer with reinforcement learning, while freezing all base weights, matches the accuracy of fully RL-tuned reasoning models on mathematical-reasoning tasks. On an 8 billion-parameter model this adds only $\approx 0.0016\%$ additional parameters and reproduces performance across a range of base models and mathematical-reasoning benchmarks. These results tighten the upper bound on the parameter budget required for high-level chain-of-thought reasoning, indicating that millions of adapter weights are unnecessary. The minimal trainable footprint reduces optimizer memory and inter-GPU communication, lowering the overall cost of fine-tuning. Moreover, a logit-lens analysis shows that the learned vectors amplify coherent token directions, providing clearer insight into the model's internal computations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_18706 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Steering LLM Reasoning Through Bias-Only Adaptation Sinii, Viacheslav Gorbatovski, Alexey Cherepanov, Artem Shaposhnikov, Boris Balagansky, Nikita Gavrilov, Daniil Machine Learning Artificial Intelligence We show that training a single $d$-dimensional steering vector per layer with reinforcement learning, while freezing all base weights, matches the accuracy of fully RL-tuned reasoning models on mathematical-reasoning tasks. On an 8 billion-parameter model this adds only $\approx 0.0016\%$ additional parameters and reproduces performance across a range of base models and mathematical-reasoning benchmarks. These results tighten the upper bound on the parameter budget required for high-level chain-of-thought reasoning, indicating that millions of adapter weights are unnecessary. The minimal trainable footprint reduces optimizer memory and inter-GPU communication, lowering the overall cost of fine-tuning. Moreover, a logit-lens analysis shows that the learned vectors amplify coherent token directions, providing clearer insight into the model's internal computations. |
| title | Steering LLM Reasoning Through Bias-Only Adaptation |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2505.18706 |