Steering Large Language Model Activations in Sparse Spaces
Fuente:
arXiv
Salvato in:
| Autori principali: | Bayat, Reza, Rahimi-Kalahroudi, Ali, Pezeshki, Mohammad, Chandar, Sarath, Vincent, Pascal |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Pitfalls of Memorization: When Memorization Hurts Generalization
di: Bayat, Reza, et al.
Pubblicazione: (2024)
di: Bayat, Reza, et al.
Pubblicazione: (2024)
Partial Models for Building Adaptive Model-Based Reinforcement Learning Agents
di: Alver, Safa, et al.
Pubblicazione: (2024)
di: Alver, Safa, et al.
Pubblicazione: (2024)
Are self-explanations from Large Language Models faithful?
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
Too Big to Fool: Resisting Deception in Language Models
di: Samsami, Mohammad Reza, et al.
Pubblicazione: (2024)
di: Samsami, Mohammad Reza, et al.
Pubblicazione: (2024)
Do Large Language Models Know How Much They Know?
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
Shielded Controller Units for RL with Operational Constraints Applied to Remote Microgrids
di: Nekoei, Hadi, et al.
Pubblicazione: (2025)
di: Nekoei, Hadi, et al.
Pubblicazione: (2025)
Compositional Risk Minimization
di: Mahajan, Divyat, et al.
Pubblicazione: (2024)
di: Mahajan, Divyat, et al.
Pubblicazione: (2024)
Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
di: Guiroy, Simon, et al.
Pubblicazione: (2025)
di: Guiroy, Simon, et al.
Pubblicazione: (2025)
Intelligent Switching for Reset-Free RL
di: Patil, Darshan, et al.
Pubblicazione: (2024)
di: Patil, Darshan, et al.
Pubblicazione: (2024)
Lookbehind-SAM: k steps back, 1 step forward
di: Mordido, Gonçalo, et al.
Pubblicazione: (2023)
di: Mordido, Gonçalo, et al.
Pubblicazione: (2023)
Amnesia: Adversarial Semantic Layer Specific Activation Steering in Large Language Models
di: Raza, Ali, et al.
Pubblicazione: (2026)
di: Raza, Ali, et al.
Pubblicazione: (2026)
Discovering environments with XRM
di: Pezeshki, Mohammad, et al.
Pubblicazione: (2023)
di: Pezeshki, Mohammad, et al.
Pubblicazione: (2023)
Squeezing More from the Stream : Learning Representation Online for Streaming Reinforcement Learning
di: Nilaksh, et al.
Pubblicazione: (2026)
di: Nilaksh, et al.
Pubblicazione: (2026)
Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models
di: Wu, Jialin, et al.
Pubblicazione: (2026)
di: Wu, Jialin, et al.
Pubblicazione: (2026)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
di: Parthasarathi, Prasanna, et al.
Pubblicazione: (2025)
di: Parthasarathi, Prasanna, et al.
Pubblicazione: (2025)
Towards Practical Tool Usage for Continually Learning LLMs
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
di: Abbes, Istabrak, et al.
Pubblicazione: (2025)
di: Abbes, Istabrak, et al.
Pubblicazione: (2025)
MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
di: Ding, Chenlu, et al.
Pubblicazione: (2025)
di: Ding, Chenlu, et al.
Pubblicazione: (2025)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
di: Soo, Samuel, et al.
Pubblicazione: (2025)
di: Soo, Samuel, et al.
Pubblicazione: (2025)
Multi-property Steering of Large Language Models with Dynamic Activation Composition
di: Scalena, Daniel, et al.
Pubblicazione: (2024)
di: Scalena, Daniel, et al.
Pubblicazione: (2024)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
di: Anand, Nikhil, et al.
Pubblicazione: (2026)
di: Anand, Nikhil, et al.
Pubblicazione: (2026)
Torque-Aware Momentum
di: Malviya, Pranshu, et al.
Pubblicazione: (2024)
di: Malviya, Pranshu, et al.
Pubblicazione: (2024)
Performance Control in Early Exiting to Deploy Large Models at the Same Cost of Smaller Ones
di: Mofakhami, Mehrnaz, et al.
Pubblicazione: (2024)
di: Mofakhami, Mehrnaz, et al.
Pubblicazione: (2024)
Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models
di: Terzić, Aleksandar, et al.
Pubblicazione: (2025)
di: Terzić, Aleksandar, et al.
Pubblicazione: (2025)
Angular Steering: Behavior Control via Rotation in Activation Space
di: Vu, Hieu M., et al.
Pubblicazione: (2025)
di: Vu, Hieu M., et al.
Pubblicazione: (2025)
Endogenous Resistance to Activation Steering in Language Models
di: McKenzie, Alex, et al.
Pubblicazione: (2026)
di: McKenzie, Alex, et al.
Pubblicazione: (2026)
SAKE: Steering Activations for Knowledge Editing
di: Scialanga, Marco, et al.
Pubblicazione: (2025)
di: Scialanga, Marco, et al.
Pubblicazione: (2025)
Promoting Exploration in Memory-Augmented Adam using Critical Momenta
di: Malviya, Pranshu, et al.
Pubblicazione: (2023)
di: Malviya, Pranshu, et al.
Pubblicazione: (2023)
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
di: Terzić, Aleksandar, et al.
Pubblicazione: (2026)
di: Terzić, Aleksandar, et al.
Pubblicazione: (2026)
Why Don't Prompt-Based Fairness Metrics Correlate?
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2023)
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2023)
Compositional Steering of Large Language Models with Steering Tokens
di: Radevski, Gorjan, et al.
Pubblicazione: (2026)
di: Radevski, Gorjan, et al.
Pubblicazione: (2026)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
di: He, Zirui, et al.
Pubblicazione: (2025)
di: He, Zirui, et al.
Pubblicazione: (2025)
On the Non-Identifiability of Steering Vectors in Large Language Models
di: Venkatesh, Sohan, et al.
Pubblicazione: (2026)
di: Venkatesh, Sohan, et al.
Pubblicazione: (2026)
Iterative Amortized Inference: Unifying In-Context Learning and Learned Optimizers
di: Mittal, Sarthak, et al.
Pubblicazione: (2025)
di: Mittal, Sarthak, et al.
Pubblicazione: (2025)
Dynamically Scaled Activation Steering
di: Ferrando, Alex, et al.
Pubblicazione: (2025)
di: Ferrando, Alex, et al.
Pubblicazione: (2025)
Uncertainty-Aware Generative Oversampling Using an Entropy-Guided Conditional Variational Autoencoder
di: Zare, Amirhossein, et al.
Pubblicazione: (2025)
di: Zare, Amirhossein, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Pitfalls of Memorization: When Memorization Hurts Generalization
di: Bayat, Reza, et al.
Pubblicazione: (2024) -
Partial Models for Building Adaptive Model-Based Reinforcement Learning Agents
di: Alver, Safa, et al.
Pubblicazione: (2024) -
Are self-explanations from Large Language Models faithful?
di: Madsen, Andreas, et al.
Pubblicazione: (2024) -
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
di: Huang, Jerry, et al.
Pubblicazione: (2024) -
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
di: Prato, Gabriele, et al.
Pubblicazione: (2025)