Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
Fuente:
arXiv
Saved in:
| Main Authors: | Sinii, Viacheslav, Balagansky, Nikita, Gerasimov, Gleb, Laptev, Daniil, Aksenov, Yaroslav, Kurochkin, Vadim, Gorbatovski, Alexey, Shaposhnikov, Boris, Gavrilov, Daniil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Steering LLM Reasoning Through Bias-Only Adaptation
by: Sinii, Viacheslav, et al.
Published: (2025)
by: Sinii, Viacheslav, et al.
Published: (2025)
You Do Not Fully Utilize Transformer's Representation Capacity
by: Gerasimov, Gleb, et al.
Published: (2025)
by: Gerasimov, Gleb, et al.
Published: (2025)
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
by: Balagansky, Nikita, et al.
Published: (2025)
by: Balagansky, Nikita, et al.
Published: (2025)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
by: Kurochkin, Vadim, et al.
Published: (2025)
by: Kurochkin, Vadim, et al.
Published: (2025)
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
by: Laptev, Daniil, et al.
Published: (2025)
by: Laptev, Daniil, et al.
Published: (2025)
Teach Old SAEs New Domain Tricks with Boosting
by: Koriagin, Nikita, et al.
Published: (2025)
by: Koriagin, Nikita, et al.
Published: (2025)
The Differences Between Direct Alignment Algorithms are a Blur
by: Gorbatovski, Alexey, et al.
Published: (2025)
by: Gorbatovski, Alexey, et al.
Published: (2025)
Learn Your Reference Model for Real Good Alignment
by: Gorbatovski, Alexey, et al.
Published: (2024)
by: Gorbatovski, Alexey, et al.
Published: (2024)
Linear Transformers with Learnable Kernel Functions are Better In-Context Models
by: Aksenov, Yaroslav, et al.
Published: (2024)
by: Aksenov, Yaroslav, et al.
Published: (2024)
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
Trust-Region Behavior Blending for On-Policy Distillation
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
Mechanistic Permutability: Match Features Across Layers
by: Balagansky, Nikita, et al.
Published: (2024)
by: Balagansky, Nikita, et al.
Published: (2024)
ESSA: Evolutionary Strategies for Scalable Alignment
by: Korotyshova, Daria, et al.
Published: (2025)
by: Korotyshova, Daria, et al.
Published: (2025)
Next Embedding Prediction Makes World Models Stronger
by: Bredis, George, et al.
Published: (2026)
by: Bredis, George, et al.
Published: (2026)
Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
by: Bredis, George, et al.
Published: (2025)
by: Bredis, George, et al.
Published: (2025)
Diffusion Language Models Generation Can Be Halted Early
by: Vaina, Sofia Maria Lo Cicero, et al.
Published: (2023)
by: Vaina, Sofia Maria Lo Cicero, et al.
Published: (2023)
Guided Star-Shaped Masked Diffusion
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
Generative Flow Networks as Entropy-Regularized RL
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
Accelerating Transformers in Online RL
by: Zelezetsky, Daniil, et al.
Published: (2025)
by: Zelezetsky, Daniil, et al.
Published: (2025)
Ensemble-based graph representation of fMRI data for cognitive brain state classification
by: Vlasenko, Daniil, et al.
Published: (2025)
by: Vlasenko, Daniil, et al.
Published: (2025)
N-Gram Induction Heads for In-Context RL: Improving Stability and Reducing Data Needs
by: Zisman, Ilya, et al.
Published: (2024)
by: Zisman, Ilya, et al.
Published: (2024)
Reinforcement learning for question answering in programming domain using public community scoring as a human feedback
by: Gorbatovski, Alexey, et al.
Published: (2024)
by: Gorbatovski, Alexey, et al.
Published: (2024)
Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages
by: Gurgurov, Daniil, et al.
Published: (2025)
by: Gurgurov, Daniil, et al.
Published: (2025)
Variance Reduction Methods Do Not Need to Compute Full Gradients: Improved Efficiency through Shuffling
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Effective Method with Compression for Distributed and Federated Cocoercive Variational Inequalities
by: Medyakov, Daniil, et al.
Published: (2024)
by: Medyakov, Daniil, et al.
Published: (2024)
Improving GFlowNets with Monte Carlo Tree Search
by: Morozov, Nikita, et al.
Published: (2024)
by: Morozov, Nikita, et al.
Published: (2024)
Improved Probabilistic Lower Bounds for Separable Matrices
by: Goshkoder, Daniil, et al.
Published: (2024)
by: Goshkoder, Daniil, et al.
Published: (2024)
Understanding Reasoning in Thinking Language Models via Steering Vectors
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Reasoning About Probabilities, Actions, and Knowledge in Fuzzy Modal Logic
by: Kozhemiachenko, Daniil, et al.
Published: (2026)
by: Kozhemiachenko, Daniil, et al.
Published: (2026)
Demonstration-Regularized RL
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
Dermacentor reticulatus genome and annotation
by: Kostarnoy, Alexey, et al.
Published: (2026)
by: Kostarnoy, Alexey, et al.
Published: (2026)
Uncertainty Estimation of Transformers' Predictions via Topological Analysis of the Attention Matrices
by: Kostenok, Elizaveta, et al.
Published: (2023)
by: Kostenok, Elizaveta, et al.
Published: (2023)
Depth scaling of unstructured search via quantum approximate optimization
by: Campos, Ernesto, et al.
Published: (2024)
by: Campos, Ernesto, et al.
Published: (2024)
Scaling Law for Sequence-Induced Demixing of Compositionally Identical Copolymers
by: Rumyantsev, Artem M., et al.
Published: (2026)
by: Rumyantsev, Artem M., et al.
Published: (2026)
Intrinsic Gaussian Vector Fields on Manifolds
by: Robert-Nicoud, Daniel, et al.
Published: (2023)
by: Robert-Nicoud, Daniel, et al.
Published: (2023)
The Euler characteristic of a triangulated manifold in terms of even-dimensional faces
by: Gavrilov, Alexey V.
Published: (2025)
by: Gavrilov, Alexey V.
Published: (2025)
Persona Vectors in Games: Measuring and Steering Strategies via Activation Vectors
by: Sun, Johnathan, et al.
Published: (2026)
by: Sun, Johnathan, et al.
Published: (2026)
Subliminal Learning Is Steering Vector Distillation
by: Blank, Camila, et al.
Published: (2026)
by: Blank, Camila, et al.
Published: (2026)
Analysing the Safety Pitfalls of Steering Vectors
by: Li, Yuxiao, et al.
Published: (2026)
by: Li, Yuxiao, et al.
Published: (2026)
Analyzing the Generalization and Reliability of Steering Vectors
by: Tan, Daniel, et al.
Published: (2024)
by: Tan, Daniel, et al.
Published: (2024)
Similar Items
-
Steering LLM Reasoning Through Bias-Only Adaptation
by: Sinii, Viacheslav, et al.
Published: (2025) -
You Do Not Fully Utilize Transformer's Representation Capacity
by: Gerasimov, Gleb, et al.
Published: (2025) -
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
by: Balagansky, Nikita, et al.
Published: (2025) -
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
by: Kurochkin, Vadim, et al.
Published: (2025) -
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
by: Laptev, Daniil, et al.
Published: (2025)