Understanding (Un)Reliability of Steering Vectors in Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Braun, Joschka, Eickhoff, Carsten, Krueger, David, Bahrainian, Seyed Ali, Krasheninnikov, Dmitrii |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
von: Braun, Joschka, et al.
Veröffentlicht: (2025)
von: Braun, Joschka, et al.
Veröffentlicht: (2025)
Logit Reweighting for Topic-Focused Summarization
von: Braun, Joschka, et al.
Veröffentlicht: (2025)
von: Braun, Joschka, et al.
Veröffentlicht: (2025)
Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
von: Braun, Joschka
Veröffentlicht: (2026)
von: Braun, Joschka
Veröffentlicht: (2026)
Comparing Bottom-Up and Top-Down Steering Approaches on In-Context Learning Tasks
von: Brumley, Madeline, et al.
Veröffentlicht: (2024)
von: Brumley, Madeline, et al.
Veröffentlicht: (2024)
Stress-Testing Capability Elicitation With Password-Locked Models
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
Implicit meta-learning may lead language models to trust more reliable sources
von: Krasheninnikov, Dmitrii, et al.
Veröffentlicht: (2023)
von: Krasheninnikov, Dmitrii, et al.
Veröffentlicht: (2023)
Fresh in memory: Training-order recency is linearly encoded in language model activations
von: Krasheninnikov, Dmitrii, et al.
Veröffentlicht: (2025)
von: Krasheninnikov, Dmitrii, et al.
Veröffentlicht: (2025)
Language Models Implement Simple Word2Vec-style Vector Arithmetic
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
Defining and Characterizing Reward Hacking
von: Skalse, Joar, et al.
Veröffentlicht: (2022)
von: Skalse, Joar, et al.
Veröffentlicht: (2022)
Analyzing the Generalization and Reliability of Steering Vectors
von: Tan, Daniel, et al.
Veröffentlicht: (2024)
von: Tan, Daniel, et al.
Veröffentlicht: (2024)
Enhancing Retrieval-Augmented Generation: A Study of Best Practices
von: Li, Siran, et al.
Veröffentlicht: (2025)
von: Li, Siran, et al.
Veröffentlicht: (2025)
Circuit Component Reuse Across Tasks in Transformer Language Models
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
Understanding Reasoning in Thinking Language Models via Steering Vectors
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
Detecting High-Stakes Interactions with Activation Probes
von: McKenzie, Alex, et al.
Veröffentlicht: (2025)
von: McKenzie, Alex, et al.
Veröffentlicht: (2025)
When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
Stable Anisotropic Regularization
von: Rudman, William, et al.
Veröffentlicht: (2023)
von: Rudman, William, et al.
Veröffentlicht: (2023)
Fill in the Blanks: Accelerating Q-Learning with a Handful of Demonstrations in Sparse Reward Settings
von: Azad, Seyed Mahdi Basiri, et al.
Veröffentlicht: (2025)
von: Azad, Seyed Mahdi Basiri, et al.
Veröffentlicht: (2025)
On the Non-Identifiability of Steering Vectors in Large Language Models
von: Venkatesh, Sohan, et al.
Veröffentlicht: (2026)
von: Venkatesh, Sohan, et al.
Veröffentlicht: (2026)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
PiCME: Pipeline for Contrastive Modality Evaluation and Encoding in the MIMIC Dataset
von: Golovanevsky, Michal, et al.
Veröffentlicht: (2025)
von: Golovanevsky, Michal, et al.
Veröffentlicht: (2025)
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
von: Cao, Yuanpu, et al.
Veröffentlicht: (2024)
von: Cao, Yuanpu, et al.
Veröffentlicht: (2024)
Benchmarking is Broken -- Don't Let AI be its Own Judge
von: Cheng, Zerui, et al.
Veröffentlicht: (2025)
von: Cheng, Zerui, et al.
Veröffentlicht: (2025)
Pixels Versus Priors: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfacts
von: Golovanevsky, Michal, et al.
Veröffentlicht: (2025)
von: Golovanevsky, Michal, et al.
Veröffentlicht: (2025)
White-Box Sensitivity Auditing with Steering Vectors
von: Cyberey, Hannah, et al.
Veröffentlicht: (2026)
von: Cyberey, Hannah, et al.
Veröffentlicht: (2026)
Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only Interventions
von: Bao, Yuntai, et al.
Veröffentlicht: (2026)
von: Bao, Yuntai, et al.
Veröffentlicht: (2026)
Predicting Where Steering Vectors Succeed
von: Billa, Jayadev
Veröffentlicht: (2026)
von: Billa, Jayadev
Veröffentlicht: (2026)
Steering Language Model Refusal with Sparse Autoencoders
von: O'Brien, Kyle, et al.
Veröffentlicht: (2024)
von: O'Brien, Kyle, et al.
Veröffentlicht: (2024)
SR-Reward: Taking The Path More Traveled
von: Azad, Seyed Mahdi B., et al.
Veröffentlicht: (2025)
von: Azad, Seyed Mahdi B., et al.
Veröffentlicht: (2025)
Steering Large Language Model Activations in Sparse Spaces
von: Bayat, Reza, et al.
Veröffentlicht: (2025)
von: Bayat, Reza, et al.
Veröffentlicht: (2025)
Explainable AI (XAI) for Arrhythmia detection from electrocardiograms
von: Beck, Joschka, et al.
Veröffentlicht: (2025)
von: Beck, Joschka, et al.
Veröffentlicht: (2025)
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
von: Jin, Zehao, et al.
Veröffentlicht: (2026)
von: Jin, Zehao, et al.
Veröffentlicht: (2026)
Steering Protein Language Models
von: Huang, Long-Kai, et al.
Veröffentlicht: (2025)
von: Huang, Long-Kai, et al.
Veröffentlicht: (2025)
Steering Language Models With Activation Engineering
von: Turner, Alexander Matt, et al.
Veröffentlicht: (2023)
von: Turner, Alexander Matt, et al.
Veröffentlicht: (2023)
One-Versus-Others Attention: Scalable Multimodal Integration for Biomedical Data
von: Golovanevsky, Michal, et al.
Veröffentlicht: (2023)
von: Golovanevsky, Michal, et al.
Veröffentlicht: (2023)
Does TabPFN Understand Causal Structures?
von: Swelam, Omar, et al.
Veröffentlicht: (2025)
von: Swelam, Omar, et al.
Veröffentlicht: (2025)
Dialz: A Python Toolkit for Steering Vectors
von: Siddique, Zara, et al.
Veröffentlicht: (2025)
von: Siddique, Zara, et al.
Veröffentlicht: (2025)
Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
von: Hedström, Anna, et al.
Veröffentlicht: (2025)
von: Hedström, Anna, et al.
Veröffentlicht: (2025)
Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
von: Sinii, Viacheslav, et al.
Veröffentlicht: (2025)
von: Sinii, Viacheslav, et al.
Veröffentlicht: (2025)
TRIM: Achieving Extreme Sparsity with Targeted Row-wise Iterative Metric-driven Pruning
von: Beck, Florentin, et al.
Veröffentlicht: (2025)
von: Beck, Florentin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
von: Braun, Joschka, et al.
Veröffentlicht: (2025) -
Logit Reweighting for Topic-Focused Summarization
von: Braun, Joschka, et al.
Veröffentlicht: (2025) -
Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
von: Braun, Joschka
Veröffentlicht: (2026) -
Comparing Bottom-Up and Top-Down Steering Approaches on In-Context Learning Tasks
von: Brumley, Madeline, et al.
Veröffentlicht: (2024) -
Stress-Testing Capability Elicitation With Password-Locked Models
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)