Towards Understanding Steering Strength
Fuente:
arXiv
Salvato in:
| Autori principali: | Taimeskhanov, Magamed, Vaiter, Samuel, Garreau, Damien |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Feature Attribution from First Principles
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2025)
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2025)
CAM-Based Methods Can See through Walls
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2024)
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2024)
MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations
di: Mitsuzawa, Kensuke, et al.
Pubblicazione: (2025)
di: Mitsuzawa, Kensuke, et al.
Pubblicazione: (2025)
Understanding Post-hoc Explainers: The Case of Anchors
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2023)
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2023)
Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022)
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022)
Faithful and Robust Local Interpretability for Textual Predictions
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2023)
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2023)
Attention Meets Post-hoc Interpretability: A Mathematical Perspective
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2024)
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2024)
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022)
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022)
A Unified Understanding and Evaluation of Steering Methods
di: Im, Shawn, et al.
Pubblicazione: (2025)
di: Im, Shawn, et al.
Pubblicazione: (2025)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
di: Lucchetti, Francesca, et al.
Pubblicazione: (2024)
di: Lucchetti, Francesca, et al.
Pubblicazione: (2024)
Understanding and Mitigating Dataset Corruption in LLM Steering
di: Anderson, Cullen, et al.
Pubblicazione: (2026)
di: Anderson, Cullen, et al.
Pubblicazione: (2026)
SteerConf: Steering LLMs for Confidence Elicitation
di: Zhou, Ziang, et al.
Pubblicazione: (2025)
di: Zhou, Ziang, et al.
Pubblicazione: (2025)
Steering Without Side Effects: Improving Post-Deployment Control of Language Models
di: Stickland, Asa Cooper, et al.
Pubblicazione: (2024)
di: Stickland, Asa Cooper, et al.
Pubblicazione: (2024)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
di: Zhang, Zhenyu, et al.
Pubblicazione: (2025)
di: Zhang, Zhenyu, et al.
Pubblicazione: (2025)
Conceptors for Semantic Steering
di: Triantafyllopoulos, Ilias, et al.
Pubblicazione: (2026)
di: Triantafyllopoulos, Ilias, et al.
Pubblicazione: (2026)
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
di: Jin, Zehao, et al.
Pubblicazione: (2026)
di: Jin, Zehao, et al.
Pubblicazione: (2026)
The Risks of Recourse in Binary Classification
di: Fokkema, Hidde, et al.
Pubblicazione: (2023)
di: Fokkema, Hidde, et al.
Pubblicazione: (2023)
That's Deprecated! Understanding, Detecting, and Steering Knowledge Conflicts in Language Models for Code Generation
di: Bae, Jaesung, et al.
Pubblicazione: (2025)
di: Bae, Jaesung, et al.
Pubblicazione: (2025)
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
di: Cao, Yuanpu, et al.
Pubblicazione: (2024)
di: Cao, Yuanpu, et al.
Pubblicazione: (2024)
On The Variability of Concept Activation Vectors
di: Wenkmann, Julia, et al.
Pubblicazione: (2025)
di: Wenkmann, Julia, et al.
Pubblicazione: (2025)
Predicting Where Steering Vectors Succeed
di: Billa, Jayadev
Pubblicazione: (2026)
di: Billa, Jayadev
Pubblicazione: (2026)
Steering Language Models with Weight Arithmetic
di: Fierro, Constanza, et al.
Pubblicazione: (2025)
di: Fierro, Constanza, et al.
Pubblicazione: (2025)
Steering Language Models With Activation Engineering
di: Turner, Alexander Matt, et al.
Pubblicazione: (2023)
di: Turner, Alexander Matt, et al.
Pubblicazione: (2023)
Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
di: Braun, Joschka
Pubblicazione: (2026)
di: Braun, Joschka
Pubblicazione: (2026)
HyperSteer: Activation Steering at Scale with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
di: Casademunt, Helena, et al.
Pubblicazione: (2025)
di: Casademunt, Helena, et al.
Pubblicazione: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
di: Heyman, Geert, et al.
Pubblicazione: (2026)
di: Heyman, Geert, et al.
Pubblicazione: (2026)
Compositional Steering of Large Language Models with Steering Tokens
di: Radevski, Gorjan, et al.
Pubblicazione: (2026)
di: Radevski, Gorjan, et al.
Pubblicazione: (2026)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
di: Soo, Samuel, et al.
Pubblicazione: (2025)
di: Soo, Samuel, et al.
Pubblicazione: (2025)
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
di: Shojaee, Parshin, et al.
Pubblicazione: (2025)
di: Shojaee, Parshin, et al.
Pubblicazione: (2025)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
di: Gan, Woody Haosheng, et al.
Pubblicazione: (2025)
di: Gan, Woody Haosheng, et al.
Pubblicazione: (2025)
VSPO: Vector-Steered Policy Optimization for Behavioral Control
di: Zhang, Xuechen, et al.
Pubblicazione: (2026)
di: Zhang, Xuechen, et al.
Pubblicazione: (2026)
ROAST: Rollout-based On-distribution Activation Steering Technique
di: Su, Xuanbo, et al.
Pubblicazione: (2026)
di: Su, Xuanbo, et al.
Pubblicazione: (2026)
Differentially Private Steering for Large Language Model Alignment
di: Goel, Anmol, et al.
Pubblicazione: (2025)
di: Goel, Anmol, et al.
Pubblicazione: (2025)
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
di: Braun, Joschka, et al.
Pubblicazione: (2025)
di: Braun, Joschka, et al.
Pubblicazione: (2025)
Towards Understanding the Relationship between In-context Learning and Compositional Generalization
di: Han, Sungjun, et al.
Pubblicazione: (2024)
di: Han, Sungjun, et al.
Pubblicazione: (2024)
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
di: He, Zhengfu, et al.
Pubblicazione: (2025)
di: He, Zhengfu, et al.
Pubblicazione: (2025)
Towards Analyzing and Understanding the Limitations of VAPO: A Theoretical Perspective
di: Shao, Jintian, et al.
Pubblicazione: (2025)
di: Shao, Jintian, et al.
Pubblicazione: (2025)
On the Hardness of Junking LLMs
di: Rando, Marco, et al.
Pubblicazione: (2026)
di: Rando, Marco, et al.
Pubblicazione: (2026)
Spherical Steering: Geometry-Aware Activation Rotation for Language Models
di: You, Zejia, et al.
Pubblicazione: (2026)
di: You, Zejia, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Feature Attribution from First Principles
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2025) -
CAM-Based Methods Can See through Walls
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2024) -
MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations
di: Mitsuzawa, Kensuke, et al.
Pubblicazione: (2025) -
Understanding Post-hoc Explainers: The Case of Anchors
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2023) -
Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022)