Steering Large Language Model Activations in Sparse Spaces
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Bayat, Reza, Rahimi-Kalahroudi, Ali, Pezeshki, Mohammad, Chandar, Sarath, Vincent, Pascal |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Pitfalls of Memorization: When Memorization Hurts Generalization
par: Bayat, Reza, et autres
Publié: (2024)
par: Bayat, Reza, et autres
Publié: (2024)
Partial Models for Building Adaptive Model-Based Reinforcement Learning Agents
par: Alver, Safa, et autres
Publié: (2024)
par: Alver, Safa, et autres
Publié: (2024)
Are self-explanations from Large Language Models faithful?
par: Madsen, Andreas, et autres
Publié: (2024)
par: Madsen, Andreas, et autres
Publié: (2024)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
par: Huang, Jerry, et autres
Publié: (2024)
par: Huang, Jerry, et autres
Publié: (2024)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
par: Prato, Gabriele, et autres
Publié: (2025)
par: Prato, Gabriele, et autres
Publié: (2025)
Too Big to Fool: Resisting Deception in Language Models
par: Samsami, Mohammad Reza, et autres
Publié: (2024)
par: Samsami, Mohammad Reza, et autres
Publié: (2024)
Do Large Language Models Know How Much They Know?
par: Prato, Gabriele, et autres
Publié: (2025)
par: Prato, Gabriele, et autres
Publié: (2025)
Shielded Controller Units for RL with Operational Constraints Applied to Remote Microgrids
par: Nekoei, Hadi, et autres
Publié: (2025)
par: Nekoei, Hadi, et autres
Publié: (2025)
Compositional Risk Minimization
par: Mahajan, Divyat, et autres
Publié: (2024)
par: Mahajan, Divyat, et autres
Publié: (2024)
Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
par: Guiroy, Simon, et autres
Publié: (2025)
par: Guiroy, Simon, et autres
Publié: (2025)
Intelligent Switching for Reset-Free RL
par: Patil, Darshan, et autres
Publié: (2024)
par: Patil, Darshan, et autres
Publié: (2024)
Lookbehind-SAM: k steps back, 1 step forward
par: Mordido, Gonçalo, et autres
Publié: (2023)
par: Mordido, Gonçalo, et autres
Publié: (2023)
Amnesia: Adversarial Semantic Layer Specific Activation Steering in Large Language Models
par: Raza, Ali, et autres
Publié: (2026)
par: Raza, Ali, et autres
Publié: (2026)
Discovering environments with XRM
par: Pezeshki, Mohammad, et autres
Publié: (2023)
par: Pezeshki, Mohammad, et autres
Publié: (2023)
Squeezing More from the Stream : Learning Representation Online for Streaming Reinforcement Learning
par: Nilaksh, et autres
Publié: (2026)
par: Nilaksh, et autres
Publié: (2026)
Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models
par: Wu, Jialin, et autres
Publié: (2026)
par: Wu, Jialin, et autres
Publié: (2026)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
par: Parthasarathi, Prasanna, et autres
Publié: (2025)
par: Parthasarathi, Prasanna, et autres
Publié: (2025)
Towards Practical Tool Usage for Continually Learning LLMs
par: Huang, Jerry, et autres
Publié: (2024)
par: Huang, Jerry, et autres
Publié: (2024)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
par: Abbes, Istabrak, et autres
Publié: (2025)
par: Abbes, Istabrak, et autres
Publié: (2025)
MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
par: Ding, Chenlu, et autres
Publié: (2025)
par: Ding, Chenlu, et autres
Publié: (2025)
Torque-Aware Momentum
par: Malviya, Pranshu, et autres
Publié: (2024)
par: Malviya, Pranshu, et autres
Publié: (2024)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
par: Soo, Samuel, et autres
Publié: (2025)
par: Soo, Samuel, et autres
Publié: (2025)
Multi-property Steering of Large Language Models with Dynamic Activation Composition
par: Scalena, Daniel, et autres
Publié: (2024)
par: Scalena, Daniel, et autres
Publié: (2024)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
par: Anand, Nikhil, et autres
Publié: (2026)
par: Anand, Nikhil, et autres
Publié: (2026)
Performance Control in Early Exiting to Deploy Large Models at the Same Cost of Smaller Ones
par: Mofakhami, Mehrnaz, et autres
Publié: (2024)
par: Mofakhami, Mehrnaz, et autres
Publié: (2024)
Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models
par: Terzić, Aleksandar, et autres
Publié: (2025)
par: Terzić, Aleksandar, et autres
Publié: (2025)
Angular Steering: Behavior Control via Rotation in Activation Space
par: Vu, Hieu M., et autres
Publié: (2025)
par: Vu, Hieu M., et autres
Publié: (2025)
Endogenous Resistance to Activation Steering in Language Models
par: McKenzie, Alex, et autres
Publié: (2026)
par: McKenzie, Alex, et autres
Publié: (2026)
SAKE: Steering Activations for Knowledge Editing
par: Scialanga, Marco, et autres
Publié: (2025)
par: Scialanga, Marco, et autres
Publié: (2025)
Promoting Exploration in Memory-Augmented Adam using Critical Momenta
par: Malviya, Pranshu, et autres
Publié: (2023)
par: Malviya, Pranshu, et autres
Publié: (2023)
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
par: Terzić, Aleksandar, et autres
Publié: (2026)
par: Terzić, Aleksandar, et autres
Publié: (2026)
Why Don't Prompt-Based Fairness Metrics Correlate?
par: Zayed, Abdelrahman, et autres
Publié: (2024)
par: Zayed, Abdelrahman, et autres
Publié: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
par: Zayed, Abdelrahman, et autres
Publié: (2023)
par: Zayed, Abdelrahman, et autres
Publié: (2023)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
par: Huang, Jerry, et autres
Publié: (2024)
par: Huang, Jerry, et autres
Publié: (2024)
Compositional Steering of Large Language Models with Steering Tokens
par: Radevski, Gorjan, et autres
Publié: (2026)
par: Radevski, Gorjan, et autres
Publié: (2026)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
par: He, Zirui, et autres
Publié: (2025)
par: He, Zirui, et autres
Publié: (2025)
Iterative Amortized Inference: Unifying In-Context Learning and Learned Optimizers
par: Mittal, Sarthak, et autres
Publié: (2025)
par: Mittal, Sarthak, et autres
Publié: (2025)
On the Non-Identifiability of Steering Vectors in Large Language Models
par: Venkatesh, Sohan, et autres
Publié: (2026)
par: Venkatesh, Sohan, et autres
Publié: (2026)
Dynamically Scaled Activation Steering
par: Ferrando, Alex, et autres
Publié: (2025)
par: Ferrando, Alex, et autres
Publié: (2025)
Uncertainty-Aware Generative Oversampling Using an Entropy-Guided Conditional Variational Autoencoder
par: Zare, Amirhossein, et autres
Publié: (2025)
par: Zare, Amirhossein, et autres
Publié: (2025)
Documents similaires
-
The Pitfalls of Memorization: When Memorization Hurts Generalization
par: Bayat, Reza, et autres
Publié: (2024) -
Partial Models for Building Adaptive Model-Based Reinforcement Learning Agents
par: Alver, Safa, et autres
Publié: (2024) -
Are self-explanations from Large Language Models faithful?
par: Madsen, Andreas, et autres
Publié: (2024) -
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
par: Huang, Jerry, et autres
Publié: (2024) -
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
par: Prato, Gabriele, et autres
Publié: (2025)