Steering Large Language Model Activations in Sparse Spaces
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bayat, Reza, Rahimi-Kalahroudi, Ali, Pezeshki, Mohammad, Chandar, Sarath, Vincent, Pascal |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Pitfalls of Memorization: When Memorization Hurts Generalization
von: Bayat, Reza, et al.
Veröffentlicht: (2024)
von: Bayat, Reza, et al.
Veröffentlicht: (2024)
Partial Models for Building Adaptive Model-Based Reinforcement Learning Agents
von: Alver, Safa, et al.
Veröffentlicht: (2024)
von: Alver, Safa, et al.
Veröffentlicht: (2024)
Are self-explanations from Large Language Models faithful?
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
Too Big to Fool: Resisting Deception in Language Models
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
Do Large Language Models Know How Much They Know?
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
Shielded Controller Units for RL with Operational Constraints Applied to Remote Microgrids
von: Nekoei, Hadi, et al.
Veröffentlicht: (2025)
von: Nekoei, Hadi, et al.
Veröffentlicht: (2025)
Compositional Risk Minimization
von: Mahajan, Divyat, et al.
Veröffentlicht: (2024)
von: Mahajan, Divyat, et al.
Veröffentlicht: (2024)
Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
von: Guiroy, Simon, et al.
Veröffentlicht: (2025)
von: Guiroy, Simon, et al.
Veröffentlicht: (2025)
Intelligent Switching for Reset-Free RL
von: Patil, Darshan, et al.
Veröffentlicht: (2024)
von: Patil, Darshan, et al.
Veröffentlicht: (2024)
Lookbehind-SAM: k steps back, 1 step forward
von: Mordido, Gonçalo, et al.
Veröffentlicht: (2023)
von: Mordido, Gonçalo, et al.
Veröffentlicht: (2023)
Amnesia: Adversarial Semantic Layer Specific Activation Steering in Large Language Models
von: Raza, Ali, et al.
Veröffentlicht: (2026)
von: Raza, Ali, et al.
Veröffentlicht: (2026)
Discovering environments with XRM
von: Pezeshki, Mohammad, et al.
Veröffentlicht: (2023)
von: Pezeshki, Mohammad, et al.
Veröffentlicht: (2023)
Squeezing More from the Stream : Learning Representation Online for Streaming Reinforcement Learning
von: Nilaksh, et al.
Veröffentlicht: (2026)
von: Nilaksh, et al.
Veröffentlicht: (2026)
Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models
von: Wu, Jialin, et al.
Veröffentlicht: (2026)
von: Wu, Jialin, et al.
Veröffentlicht: (2026)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025)
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025)
Towards Practical Tool Usage for Continually Learning LLMs
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
von: Ding, Chenlu, et al.
Veröffentlicht: (2025)
von: Ding, Chenlu, et al.
Veröffentlicht: (2025)
Torque-Aware Momentum
von: Malviya, Pranshu, et al.
Veröffentlicht: (2024)
von: Malviya, Pranshu, et al.
Veröffentlicht: (2024)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
Multi-property Steering of Large Language Models with Dynamic Activation Composition
von: Scalena, Daniel, et al.
Veröffentlicht: (2024)
von: Scalena, Daniel, et al.
Veröffentlicht: (2024)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
von: Anand, Nikhil, et al.
Veröffentlicht: (2026)
von: Anand, Nikhil, et al.
Veröffentlicht: (2026)
Performance Control in Early Exiting to Deploy Large Models at the Same Cost of Smaller Ones
von: Mofakhami, Mehrnaz, et al.
Veröffentlicht: (2024)
von: Mofakhami, Mehrnaz, et al.
Veröffentlicht: (2024)
Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2025)
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2025)
Angular Steering: Behavior Control via Rotation in Activation Space
von: Vu, Hieu M., et al.
Veröffentlicht: (2025)
von: Vu, Hieu M., et al.
Veröffentlicht: (2025)
Endogenous Resistance to Activation Steering in Language Models
von: McKenzie, Alex, et al.
Veröffentlicht: (2026)
von: McKenzie, Alex, et al.
Veröffentlicht: (2026)
SAKE: Steering Activations for Knowledge Editing
von: Scialanga, Marco, et al.
Veröffentlicht: (2025)
von: Scialanga, Marco, et al.
Veröffentlicht: (2025)
Promoting Exploration in Memory-Augmented Adam using Critical Momenta
von: Malviya, Pranshu, et al.
Veröffentlicht: (2023)
von: Malviya, Pranshu, et al.
Veröffentlicht: (2023)
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2026)
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2026)
Why Don't Prompt-Based Fairness Metrics Correlate?
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Compositional Steering of Large Language Models with Steering Tokens
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
von: He, Zirui, et al.
Veröffentlicht: (2025)
von: He, Zirui, et al.
Veröffentlicht: (2025)
Iterative Amortized Inference: Unifying In-Context Learning and Learned Optimizers
von: Mittal, Sarthak, et al.
Veröffentlicht: (2025)
von: Mittal, Sarthak, et al.
Veröffentlicht: (2025)
On the Non-Identifiability of Steering Vectors in Large Language Models
von: Venkatesh, Sohan, et al.
Veröffentlicht: (2026)
von: Venkatesh, Sohan, et al.
Veröffentlicht: (2026)
Dynamically Scaled Activation Steering
von: Ferrando, Alex, et al.
Veröffentlicht: (2025)
von: Ferrando, Alex, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Generative Oversampling Using an Entropy-Guided Conditional Variational Autoencoder
von: Zare, Amirhossein, et al.
Veröffentlicht: (2025)
von: Zare, Amirhossein, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Pitfalls of Memorization: When Memorization Hurts Generalization
von: Bayat, Reza, et al.
Veröffentlicht: (2024) -
Partial Models for Building Adaptive Model-Based Reinforcement Learning Agents
von: Alver, Safa, et al.
Veröffentlicht: (2024) -
Are self-explanations from Large Language Models faithful?
von: Madsen, Andreas, et al.
Veröffentlicht: (2024) -
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
von: Huang, Jerry, et al.
Veröffentlicht: (2024) -
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)