Towards Understanding Steering Strength
Fuente:
arXiv
Saved in:
| Main Authors: | Taimeskhanov, Magamed, Vaiter, Samuel, Garreau, Damien |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Feature Attribution from First Principles
by: Taimeskhanov, Magamed, et al.
Published: (2025)
by: Taimeskhanov, Magamed, et al.
Published: (2025)
CAM-Based Methods Can See through Walls
by: Taimeskhanov, Magamed, et al.
Published: (2024)
by: Taimeskhanov, Magamed, et al.
Published: (2024)
MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations
by: Mitsuzawa, Kensuke, et al.
Published: (2025)
by: Mitsuzawa, Kensuke, et al.
Published: (2025)
Understanding Post-hoc Explainers: The Case of Anchors
by: Lopardo, Gianluigi, et al.
Published: (2023)
by: Lopardo, Gianluigi, et al.
Published: (2023)
Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
by: Lopardo, Gianluigi, et al.
Published: (2022)
by: Lopardo, Gianluigi, et al.
Published: (2022)
Faithful and Robust Local Interpretability for Textual Predictions
by: Lopardo, Gianluigi, et al.
Published: (2023)
by: Lopardo, Gianluigi, et al.
Published: (2023)
Attention Meets Post-hoc Interpretability: A Mathematical Perspective
by: Lopardo, Gianluigi, et al.
Published: (2024)
by: Lopardo, Gianluigi, et al.
Published: (2024)
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
by: Lopardo, Gianluigi, et al.
Published: (2022)
by: Lopardo, Gianluigi, et al.
Published: (2022)
A Unified Understanding and Evaluation of Steering Methods
by: Im, Shawn, et al.
Published: (2025)
by: Im, Shawn, et al.
Published: (2025)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
by: Lucchetti, Francesca, et al.
Published: (2024)
by: Lucchetti, Francesca, et al.
Published: (2024)
Understanding and Mitigating Dataset Corruption in LLM Steering
by: Anderson, Cullen, et al.
Published: (2026)
by: Anderson, Cullen, et al.
Published: (2026)
SteerConf: Steering LLMs for Confidence Elicitation
by: Zhou, Ziang, et al.
Published: (2025)
by: Zhou, Ziang, et al.
Published: (2025)
Steering Without Side Effects: Improving Post-Deployment Control of Language Models
by: Stickland, Asa Cooper, et al.
Published: (2024)
by: Stickland, Asa Cooper, et al.
Published: (2024)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
Conceptors for Semantic Steering
by: Triantafyllopoulos, Ilias, et al.
Published: (2026)
by: Triantafyllopoulos, Ilias, et al.
Published: (2026)
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
by: Jin, Zehao, et al.
Published: (2026)
by: Jin, Zehao, et al.
Published: (2026)
The Risks of Recourse in Binary Classification
by: Fokkema, Hidde, et al.
Published: (2023)
by: Fokkema, Hidde, et al.
Published: (2023)
That's Deprecated! Understanding, Detecting, and Steering Knowledge Conflicts in Language Models for Code Generation
by: Bae, Jaesung, et al.
Published: (2025)
by: Bae, Jaesung, et al.
Published: (2025)
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
by: Cao, Yuanpu, et al.
Published: (2024)
by: Cao, Yuanpu, et al.
Published: (2024)
On The Variability of Concept Activation Vectors
by: Wenkmann, Julia, et al.
Published: (2025)
by: Wenkmann, Julia, et al.
Published: (2025)
Predicting Where Steering Vectors Succeed
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
Steering Language Models with Weight Arithmetic
by: Fierro, Constanza, et al.
Published: (2025)
by: Fierro, Constanza, et al.
Published: (2025)
Steering Language Models With Activation Engineering
by: Turner, Alexander Matt, et al.
Published: (2023)
by: Turner, Alexander Matt, et al.
Published: (2023)
Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
by: Braun, Joschka
Published: (2026)
by: Braun, Joschka
Published: (2026)
HyperSteer: Activation Steering at Scale with Hypernetworks
by: Sun, Jiuding, et al.
Published: (2025)
by: Sun, Jiuding, et al.
Published: (2025)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
by: Casademunt, Helena, et al.
Published: (2025)
by: Casademunt, Helena, et al.
Published: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
by: Heyman, Geert, et al.
Published: (2026)
by: Heyman, Geert, et al.
Published: (2026)
Compositional Steering of Large Language Models with Steering Tokens
by: Radevski, Gorjan, et al.
Published: (2026)
by: Radevski, Gorjan, et al.
Published: (2026)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
by: Soo, Samuel, et al.
Published: (2025)
by: Soo, Samuel, et al.
Published: (2025)
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
by: Shojaee, Parshin, et al.
Published: (2025)
by: Shojaee, Parshin, et al.
Published: (2025)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
by: Gan, Woody Haosheng, et al.
Published: (2025)
by: Gan, Woody Haosheng, et al.
Published: (2025)
VSPO: Vector-Steered Policy Optimization for Behavioral Control
by: Zhang, Xuechen, et al.
Published: (2026)
by: Zhang, Xuechen, et al.
Published: (2026)
ROAST: Rollout-based On-distribution Activation Steering Technique
by: Su, Xuanbo, et al.
Published: (2026)
by: Su, Xuanbo, et al.
Published: (2026)
Differentially Private Steering for Large Language Model Alignment
by: Goel, Anmol, et al.
Published: (2025)
by: Goel, Anmol, et al.
Published: (2025)
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
by: Braun, Joschka, et al.
Published: (2025)
by: Braun, Joschka, et al.
Published: (2025)
Towards Understanding the Relationship between In-context Learning and Compositional Generalization
by: Han, Sungjun, et al.
Published: (2024)
by: Han, Sungjun, et al.
Published: (2024)
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
by: He, Zhengfu, et al.
Published: (2025)
by: He, Zhengfu, et al.
Published: (2025)
Towards Analyzing and Understanding the Limitations of VAPO: A Theoretical Perspective
by: Shao, Jintian, et al.
Published: (2025)
by: Shao, Jintian, et al.
Published: (2025)
On the Hardness of Junking LLMs
by: Rando, Marco, et al.
Published: (2026)
by: Rando, Marco, et al.
Published: (2026)
Spherical Steering: Geometry-Aware Activation Rotation for Language Models
by: You, Zejia, et al.
Published: (2026)
by: You, Zejia, et al.
Published: (2026)
Similar Items
-
Feature Attribution from First Principles
by: Taimeskhanov, Magamed, et al.
Published: (2025) -
CAM-Based Methods Can See through Walls
by: Taimeskhanov, Magamed, et al.
Published: (2024) -
MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations
by: Mitsuzawa, Kensuke, et al.
Published: (2025) -
Understanding Post-hoc Explainers: The Case of Anchors
by: Lopardo, Gianluigi, et al.
Published: (2023) -
Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
by: Lopardo, Gianluigi, et al.
Published: (2022)