Predicting Where Steering Vectors Succeed
Fuente:
arXiv
Salvato in:
| Autore principale: | Billa, Jayadev |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
di: Billa, Jayadev
Pubblicazione: (2026)
di: Billa, Jayadev
Pubblicazione: (2026)
The Geometric Anatomy of Capability Acquisition in Transformers
di: Billa, Jayadev
Pubblicazione: (2026)
di: Billa, Jayadev
Pubblicazione: (2026)
Decomposing the Depth Profile of Fine-Tuning
di: Billa, Jayadev
Pubblicazione: (2026)
di: Billa, Jayadev
Pubblicazione: (2026)
When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration
di: Billa, Jayadev
Pubblicazione: (2026)
di: Billa, Jayadev
Pubblicazione: (2026)
The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
di: Billa, Jayadev
Pubblicazione: (2026)
di: Billa, Jayadev
Pubblicazione: (2026)
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
di: Jin, Zehao, et al.
Pubblicazione: (2026)
di: Jin, Zehao, et al.
Pubblicazione: (2026)
VSPO: Vector-Steered Policy Optimization for Behavioral Control
di: Zhang, Xuechen, et al.
Pubblicazione: (2026)
di: Zhang, Xuechen, et al.
Pubblicazione: (2026)
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
di: Braun, Joschka, et al.
Pubblicazione: (2025)
di: Braun, Joschka, et al.
Pubblicazione: (2025)
White-Box Sensitivity Auditing with Steering Vectors
di: Cyberey, Hannah, et al.
Pubblicazione: (2026)
di: Cyberey, Hannah, et al.
Pubblicazione: (2026)
Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs
di: Zhang, Qingru, et al.
Pubblicazione: (2023)
di: Zhang, Qingru, et al.
Pubblicazione: (2023)
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
di: Cao, Yuanpu, et al.
Pubblicazione: (2024)
di: Cao, Yuanpu, et al.
Pubblicazione: (2024)
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
di: Han, Pengrui, et al.
Pubblicazione: (2026)
di: Han, Pengrui, et al.
Pubblicazione: (2026)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)
When and Why Does Unsupervised RL Succeed in Mathematical Reasoning? A Manifold Envelopment Perspective
di: Zhang, Zelin, et al.
Pubblicazione: (2026)
di: Zhang, Zelin, et al.
Pubblicazione: (2026)
Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
di: Siddique, Zara, et al.
Pubblicazione: (2025)
di: Siddique, Zara, et al.
Pubblicazione: (2025)
No Training Wheels: Steering Vectors for Bias Correction at Inference Time
di: Gupta, Aviral, et al.
Pubblicazione: (2025)
di: Gupta, Aviral, et al.
Pubblicazione: (2025)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
di: Lucchetti, Francesca, et al.
Pubblicazione: (2024)
di: Lucchetti, Francesca, et al.
Pubblicazione: (2024)
SteerConf: Steering LLMs for Confidence Elicitation
di: Zhou, Ziang, et al.
Pubblicazione: (2025)
di: Zhou, Ziang, et al.
Pubblicazione: (2025)
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
di: Puerto, Haritz, et al.
Pubblicazione: (2024)
di: Puerto, Haritz, et al.
Pubblicazione: (2024)
Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
di: Braun, Joschka
Pubblicazione: (2026)
di: Braun, Joschka
Pubblicazione: (2026)
In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
di: Liu, Sheng, et al.
Pubblicazione: (2023)
di: Liu, Sheng, et al.
Pubblicazione: (2023)
Conceptors for Semantic Steering
di: Triantafyllopoulos, Ilias, et al.
Pubblicazione: (2026)
di: Triantafyllopoulos, Ilias, et al.
Pubblicazione: (2026)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
di: Gan, Woody Haosheng, et al.
Pubblicazione: (2025)
di: Gan, Woody Haosheng, et al.
Pubblicazione: (2025)
Behavioral Steering in a 35B MoE Language Model via SAE-Decoded Probe Vectors: One Agency Axis, Not Five Traits
di: Yap, Jia Qing
Pubblicazione: (2026)
di: Yap, Jia Qing
Pubblicazione: (2026)
Towards Understanding Steering Strength
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2026)
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2026)
Where is the signal in tokenization space?
di: Geh, Renato Lui, et al.
Pubblicazione: (2024)
di: Geh, Renato Lui, et al.
Pubblicazione: (2024)
MorphNAS: Differentiable Architecture Search for Morphologically-Aware Multilingual NER
di: Devadiga, Prathamesh, et al.
Pubblicazione: (2025)
di: Devadiga, Prathamesh, et al.
Pubblicazione: (2025)
Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute
di: Liu, Sheng, et al.
Pubblicazione: (2025)
di: Liu, Sheng, et al.
Pubblicazione: (2025)
Steering Language Models with Weight Arithmetic
di: Fierro, Constanza, et al.
Pubblicazione: (2025)
di: Fierro, Constanza, et al.
Pubblicazione: (2025)
Steering Language Models With Activation Engineering
di: Turner, Alexander Matt, et al.
Pubblicazione: (2023)
di: Turner, Alexander Matt, et al.
Pubblicazione: (2023)
HyperSteer: Activation Steering at Scale with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
A Unified Understanding and Evaluation of Steering Methods
di: Im, Shawn, et al.
Pubblicazione: (2025)
di: Im, Shawn, et al.
Pubblicazione: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
di: Heyman, Geert, et al.
Pubblicazione: (2026)
di: Heyman, Geert, et al.
Pubblicazione: (2026)
Compositional Steering of Large Language Models with Steering Tokens
di: Radevski, Gorjan, et al.
Pubblicazione: (2026)
di: Radevski, Gorjan, et al.
Pubblicazione: (2026)
ROAST: Rollout-based On-distribution Activation Steering Technique
di: Su, Xuanbo, et al.
Pubblicazione: (2026)
di: Su, Xuanbo, et al.
Pubblicazione: (2026)
Differentially Private Steering for Large Language Model Alignment
di: Goel, Anmol, et al.
Pubblicazione: (2025)
di: Goel, Anmol, et al.
Pubblicazione: (2025)
Compute Where it Counts: Self Optimizing Language Models
di: Akhauri, Yash, et al.
Pubblicazione: (2026)
di: Akhauri, Yash, et al.
Pubblicazione: (2026)
Spherical Steering: Geometry-Aware Activation Rotation for Language Models
di: You, Zejia, et al.
Pubblicazione: (2026)
di: You, Zejia, et al.
Pubblicazione: (2026)
Guiding Giants: Lightweight Controllers for Weighted Activation Steering in LLMs
di: Hegazy, Amr, et al.
Pubblicazione: (2025)
di: Hegazy, Amr, et al.
Pubblicazione: (2025)
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
di: Laptev, Daniil, et al.
Pubblicazione: (2025)
di: Laptev, Daniil, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
di: Billa, Jayadev
Pubblicazione: (2026) -
The Geometric Anatomy of Capability Acquisition in Transformers
di: Billa, Jayadev
Pubblicazione: (2026) -
Decomposing the Depth Profile of Fine-Tuning
di: Billa, Jayadev
Pubblicazione: (2026) -
When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration
di: Billa, Jayadev
Pubblicazione: (2026) -
The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
di: Billa, Jayadev
Pubblicazione: (2026)