HyperSteer: Activation Steering at Scale with Hypernetworks
Fuente:
arXiv
Guardado en:
| Autores principales: | Sun, Jiuding, Baskaran, Sidharth, Wu, Zhengxuan, Sklar, Michael, Potts, Christopher, Geiger, Atticus |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
por: Sun, Jiuding, et al.
Publicado: (2025)
por: Sun, Jiuding, et al.
Publicado: (2025)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
por: Wu, Zhengxuan, et al.
Publicado: (2025)
por: Wu, Zhengxuan, et al.
Publicado: (2025)
ReFT: Representation Finetuning for Language Models
por: Wu, Zhengxuan, et al.
Publicado: (2024)
por: Wu, Zhengxuan, et al.
Publicado: (2024)
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
por: Wu, Zhengxuan, et al.
Publicado: (2024)
por: Wu, Zhengxuan, et al.
Publicado: (2024)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
por: Huang, Jing, et al.
Publicado: (2024)
por: Huang, Jing, et al.
Publicado: (2024)
Steer Like the LLM: Activation Steering that Mimics Prompting
por: Heyman, Geert, et al.
Publicado: (2026)
por: Heyman, Geert, et al.
Publicado: (2026)
Activation Steering via Generative Causal Mediation
por: Sankaranarayanan, Aruna, et al.
Publicado: (2026)
por: Sankaranarayanan, Aruna, et al.
Publicado: (2026)
How Do Transformers Learn Variable Binding in Symbolic Programs?
por: Wu, Yiwei, et al.
Publicado: (2025)
por: Wu, Yiwei, et al.
Publicado: (2025)
Endogenous Resistance to Activation Steering in Language Models
por: McKenzie, Alex, et al.
Publicado: (2026)
por: McKenzie, Alex, et al.
Publicado: (2026)
SAKE: Steering Activations for Knowledge Editing
por: Scialanga, Marco, et al.
Publicado: (2025)
por: Scialanga, Marco, et al.
Publicado: (2025)
Programming Refusal with Conditional Activation Steering
por: Lee, Bruce W., et al.
Publicado: (2024)
por: Lee, Bruce W., et al.
Publicado: (2024)
How Causal Abstraction Underpins Computational Explanation
por: Geiger, Atticus, et al.
Publicado: (2025)
por: Geiger, Atticus, et al.
Publicado: (2025)
Extracting Unlearned Information from LLMs with Activation Steering
por: Seyitoğlu, Atakan, et al.
Publicado: (2024)
por: Seyitoğlu, Atakan, et al.
Publicado: (2024)
Steering Llama 2 via Contrastive Activation Addition
por: Panickssery, Nina, et al.
Publicado: (2023)
por: Panickssery, Nina, et al.
Publicado: (2023)
Compositional Steering of Large Language Models with Steering Tokens
por: Radevski, Gorjan, et al.
Publicado: (2026)
por: Radevski, Gorjan, et al.
Publicado: (2026)
Improving Instruction-Following in Language Models through Activation Steering
por: Stolfo, Alessandro, et al.
Publicado: (2024)
por: Stolfo, Alessandro, et al.
Publicado: (2024)
Extending Activation Steering to Broad Skills and Multiple Behaviours
por: van der Weij, Teun, et al.
Publicado: (2024)
por: van der Weij, Teun, et al.
Publicado: (2024)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
por: Bigelow, Eric, et al.
Publicado: (2025)
por: Bigelow, Eric, et al.
Publicado: (2025)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
por: Soo, Samuel, et al.
Publicado: (2025)
por: Soo, Samuel, et al.
Publicado: (2025)
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
por: Han, Pengrui, et al.
Publicado: (2026)
por: Han, Pengrui, et al.
Publicado: (2026)
Multi-property Steering of Large Language Models with Dynamic Activation Composition
por: Scalena, Daniel, et al.
Publicado: (2024)
por: Scalena, Daniel, et al.
Publicado: (2024)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
por: Anand, Nikhil, et al.
Publicado: (2026)
por: Anand, Nikhil, et al.
Publicado: (2026)
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
por: Csordás, Róbert, et al.
Publicado: (2024)
por: Csordás, Róbert, et al.
Publicado: (2024)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
por: Siu, Vincent, et al.
Publicado: (2025)
por: Siu, Vincent, et al.
Publicado: (2025)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
por: Cheng, Stephen, et al.
Publicado: (2026)
por: Cheng, Stephen, et al.
Publicado: (2026)
Word Embeddings Are Steers for Language Models
por: Han, Chi, et al.
Publicado: (2023)
por: Han, Chi, et al.
Publicado: (2023)
COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
por: Sharma, Kartik, et al.
Publicado: (2026)
por: Sharma, Kartik, et al.
Publicado: (2026)
FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models
por: Weng, Zixuan, et al.
Publicado: (2026)
por: Weng, Zixuan, et al.
Publicado: (2026)
Steering LLMs for Formal Theorem Proving
por: Kirtania, Shashank, et al.
Publicado: (2025)
por: Kirtania, Shashank, et al.
Publicado: (2025)
Steer LLM Latents for Hallucination Detection
por: Park, Seongheon, et al.
Publicado: (2025)
por: Park, Seongheon, et al.
Publicado: (2025)
The Information Geometry of Softmax: Probing and Steering
por: Park, Kiho, et al.
Publicado: (2026)
por: Park, Kiho, et al.
Publicado: (2026)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
por: Zhang, Zhenyu, et al.
Publicado: (2025)
por: Zhang, Zhenyu, et al.
Publicado: (2025)
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
por: Batra, Shourya, et al.
Publicado: (2025)
por: Batra, Shourya, et al.
Publicado: (2025)
Learning to Steer: Input-dependent Steering for Multimodal LLMs
por: Parekh, Jayneel, et al.
Publicado: (2025)
por: Parekh, Jayneel, et al.
Publicado: (2025)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
por: Lamb, Tom A., et al.
Publicado: (2024)
por: Lamb, Tom A., et al.
Publicado: (2024)
Understanding and Mitigating Dataset Corruption in LLM Steering
por: Anderson, Cullen, et al.
Publicado: (2026)
por: Anderson, Cullen, et al.
Publicado: (2026)
Dynamically Scaled Activation Steering
por: Ferrando, Alex, et al.
Publicado: (2025)
por: Ferrando, Alex, et al.
Publicado: (2025)
pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
por: Wu, Zhengxuan, et al.
Publicado: (2024)
por: Wu, Zhengxuan, et al.
Publicado: (2024)
Steering Large Language Models for Machine Translation Personalization
por: Scalena, Daniel, et al.
Publicado: (2025)
por: Scalena, Daniel, et al.
Publicado: (2025)
Brain-Grounded Axes for Reading and Steering LLM States
por: Andric, Sandro
Publicado: (2025)
por: Andric, Sandro
Publicado: (2025)
Ejemplares similares
-
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
por: Sun, Jiuding, et al.
Publicado: (2025) -
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
por: Wu, Zhengxuan, et al.
Publicado: (2025) -
ReFT: Representation Finetuning for Language Models
por: Wu, Zhengxuan, et al.
Publicado: (2024) -
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
por: Wu, Zhengxuan, et al.
Publicado: (2024) -
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
por: Huang, Jing, et al.
Publicado: (2024)