Improving Steering Vectors by Targeting Sparse Autoencoder Features
Fuente:
arXiv
Guardado en:
| Autores principales: | Chalnev, Sviatoslav, Siu, Matthew, Conmy, Arthur |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
por: Lieberum, Tom, et al.
Publicado: (2024)
por: Lieberum, Tom, et al.
Publicado: (2024)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
por: Yu, Zhuohao, et al.
Publicado: (2025)
por: Yu, Zhuohao, et al.
Publicado: (2025)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
por: Jørgensen, Mikkel Godsk, et al.
Publicado: (2026)
por: Jørgensen, Mikkel Godsk, et al.
Publicado: (2026)
Sparse Autoencoder Features for Classifications and Transferability
por: Gallifant, Jack, et al.
Publicado: (2025)
por: Gallifant, Jack, et al.
Publicado: (2025)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2025)
por: Chanin, David, et al.
Publicado: (2025)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
por: He, Zirui, et al.
Publicado: (2025)
por: He, Zirui, et al.
Publicado: (2025)
Improving Dictionary Learning with Gated Sparse Autoencoders
por: Rajamanoharan, Senthooran, et al.
Publicado: (2024)
por: Rajamanoharan, Senthooran, et al.
Publicado: (2024)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2025)
por: Chanin, David, et al.
Publicado: (2025)
UniMaia: Steering Chess Policies with Language for Human-like Play
por: Siu, Sherman, et al.
Publicado: (2026)
por: Siu, Sherman, et al.
Publicado: (2026)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
por: Hua, Zhenglin, et al.
Publicado: (2025)
por: Hua, Zhenglin, et al.
Publicado: (2025)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
por: Cho, Seonglae, et al.
Publicado: (2025)
por: Cho, Seonglae, et al.
Publicado: (2025)
Understanding Reasoning in Thinking Language Models via Steering Vectors
por: Venhoff, Constantin, et al.
Publicado: (2025)
por: Venhoff, Constantin, et al.
Publicado: (2025)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
por: Zhu, Xudong, et al.
Publicado: (2025)
por: Zhu, Xudong, et al.
Publicado: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
por: Zhao, Haiyan, et al.
Publicado: (2025)
por: Zhao, Haiyan, et al.
Publicado: (2025)
Thought Anchors: Which LLM Reasoning Steps Matter?
por: Bogdan, Paul C., et al.
Publicado: (2025)
por: Bogdan, Paul C., et al.
Publicado: (2025)
Scaling sparse feature circuit finding for in-context learning
por: Kharlapenko, Dmitrii, et al.
Publicado: (2025)
por: Kharlapenko, Dmitrii, et al.
Publicado: (2025)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
por: Lan, Michael, et al.
Publicado: (2024)
por: Lan, Michael, et al.
Publicado: (2024)
Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training
por: Li, T. Ed, et al.
Publicado: (2025)
por: Li, T. Ed, et al.
Publicado: (2025)
LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering
por: Wong, Sing Hieng, et al.
Publicado: (2026)
por: Wong, Sing Hieng, et al.
Publicado: (2026)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
por: Siu, Vincent, et al.
Publicado: (2025)
por: Siu, Vincent, et al.
Publicado: (2025)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
por: Wu, Zhengxuan, et al.
Publicado: (2025)
por: Wu, Zhengxuan, et al.
Publicado: (2025)
Dissecting Chronos: Sparse Autoencoders Reveal Causal Feature Hierarchies in Time Series Foundation Models
por: Mishra, Anurag
Publicado: (2026)
por: Mishra, Anurag
Publicado: (2026)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
por: Cho, Seonglae, et al.
Publicado: (2025)
por: Cho, Seonglae, et al.
Publicado: (2025)
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
por: Han, Pengrui, et al.
Publicado: (2026)
por: Han, Pengrui, et al.
Publicado: (2026)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
por: Arcuschin, Iván, et al.
Publicado: (2025)
por: Arcuschin, Iván, et al.
Publicado: (2025)
How do LLMs Compute Verbal Confidence
por: Kumaran, Dharshan, et al.
Publicado: (2026)
por: Kumaran, Dharshan, et al.
Publicado: (2026)
Sparse Autoencoder Decomposition of Clinical Sequence Model Representations: Feature Complexity, Task Specialisation, and Mortality Prediction
por: Sainsbury, Chris, et al.
Publicado: (2026)
por: Sainsbury, Chris, et al.
Publicado: (2026)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
por: Cho, Seonglae, et al.
Publicado: (2026)
por: Cho, Seonglae, et al.
Publicado: (2026)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
por: Muchane, Mark, et al.
Publicado: (2025)
por: Muchane, Mark, et al.
Publicado: (2025)
Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
por: Siddique, Zara, et al.
Publicado: (2025)
por: Siddique, Zara, et al.
Publicado: (2025)
Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute
por: Liu, Sheng, et al.
Publicado: (2025)
por: Liu, Sheng, et al.
Publicado: (2025)
Building Production-Ready Probes For Gemini
por: Kramár, János, et al.
Publicado: (2026)
por: Kramár, János, et al.
Publicado: (2026)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
por: Farnik, Lucy, et al.
Publicado: (2025)
por: Farnik, Lucy, et al.
Publicado: (2025)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
por: Li, Aaron J., et al.
Publicado: (2025)
por: Li, Aaron J., et al.
Publicado: (2025)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
por: Lamb, Tom A., et al.
Publicado: (2024)
por: Lamb, Tom A., et al.
Publicado: (2024)
Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
por: Wang, Anyi, et al.
Publicado: (2025)
por: Wang, Anyi, et al.
Publicado: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
por: Bhalla, Usha, et al.
Publicado: (2025)
por: Bhalla, Usha, et al.
Publicado: (2025)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
por: Minegishi, Gouki, et al.
Publicado: (2025)
por: Minegishi, Gouki, et al.
Publicado: (2025)
SAEs Are Good for Steering -- If You Select the Right Features
por: Arad, Dana, et al.
Publicado: (2025)
por: Arad, Dana, et al.
Publicado: (2025)
Multi-Attribute Steering of Language Models via Targeted Intervention
por: Nguyen, Duy, et al.
Publicado: (2025)
por: Nguyen, Duy, et al.
Publicado: (2025)
Ejemplares similares
-
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
por: Lieberum, Tom, et al.
Publicado: (2024) -
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
por: Yu, Zhuohao, et al.
Publicado: (2025) -
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
por: Jørgensen, Mikkel Godsk, et al.
Publicado: (2026) -
Sparse Autoencoder Features for Classifications and Transferability
por: Gallifant, Jack, et al.
Publicado: (2025) -
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2025)