Can sparse autoencoders be used to decompose and interpret steering vectors?
Fuente:
arXiv
Saved in:
| Main Authors: | Mayne, Harry, Yang, Yushi, Mahdi, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025)
by: Mayne, Harry, et al.
Published: (2025)
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
by: Yang, Yushi, et al.
Published: (2024)
by: Yang, Yushi, et al.
Published: (2024)
Negation Neglect: When models fail to learn negations in training
by: Mayne, Harry, et al.
Published: (2026)
by: Mayne, Harry, et al.
Published: (2026)
Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
by: Sondej, Filip, et al.
Published: (2025)
by: Sondej, Filip, et al.
Published: (2025)
Multi-view autoencoders for Fake News Detection
by: Pereira, Ingryd V. S. T., et al.
Published: (2025)
by: Pereira, Ingryd V. S. T., et al.
Published: (2025)
Toward universal steering and monitoring of AI models
by: Beaglehole, Daniel, et al.
Published: (2025)
by: Beaglehole, Daniel, et al.
Published: (2025)
Scaling and evaluating sparse autoencoders
by: Gao, Leo, et al.
Published: (2024)
by: Gao, Leo, et al.
Published: (2024)
Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
by: Sondej, Filip, et al.
Published: (2025)
by: Sondej, Filip, et al.
Published: (2025)
Efficient semantic uncertainty quantification in language models via diversity-steered sampling
by: Park, Ji Won, et al.
Published: (2025)
by: Park, Ji Won, et al.
Published: (2025)
PrivacyMind: Large Language Models Can Be Contextual Privacy Protection Learners
by: Xiao, Yijia, et al.
Published: (2023)
by: Xiao, Yijia, et al.
Published: (2023)
Steering CLIP's vision transformer with sparse autoencoders
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
Scaling sparse feature circuit finding for in-context learning
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
Agentic-imodels: Evolving agentic interpretability tools via autoresearch
by: Singh, Chandan, et al.
Published: (2026)
by: Singh, Chandan, et al.
Published: (2026)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
by: Kirchhof, Michael, et al.
Published: (2025)
by: Kirchhof, Michael, et al.
Published: (2025)
Applying sparse autoencoders to unlearn knowledge in language models
by: Farrell, Eoin, et al.
Published: (2024)
by: Farrell, Eoin, et al.
Published: (2024)
Can LLMs Reliably Simulate Human Learner Actions? A Simulation Authoring Framework for Open-Ended Learning Environments
by: Mannekote, Amogh, et al.
Published: (2024)
by: Mannekote, Amogh, et al.
Published: (2024)
LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit
by: Gong, Ruihao, et al.
Published: (2024)
by: Gong, Ruihao, et al.
Published: (2024)
From communities to interpretable network and word embedding: an unified approach
by: Prouteau, Thibault, et al.
Published: (2024)
by: Prouteau, Thibault, et al.
Published: (2024)
Surrogate modeling for interpreting black-box LLMs in medical predictions
by: Han, Changho, et al.
Published: (2026)
by: Han, Changho, et al.
Published: (2026)
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
by: Mayne, Harry, et al.
Published: (2026)
by: Mayne, Harry, et al.
Published: (2026)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
by: Chen, Jianhui, et al.
Published: (2024)
by: Chen, Jianhui, et al.
Published: (2024)
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
by: Wu, Yuhao, et al.
Published: (2025)
by: Wu, Yuhao, et al.
Published: (2025)
AcquisitionSynthesis: Targeted Data Generation using Acquisition Functions
by: Agarwal, Ishika, et al.
Published: (2026)
by: Agarwal, Ishika, et al.
Published: (2026)
ExpertLens: Activation steering features are highly interpretable
by: Fedzechkina, Masha, et al.
Published: (2025)
by: Fedzechkina, Masha, et al.
Published: (2025)
MIMIC-RD: Can LLMs differentially diagnose rare diseases in real-world clinical settings?
by: AlDin, Zilal Eiz, et al.
Published: (2025)
by: AlDin, Zilal Eiz, et al.
Published: (2025)
Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
by: Dong, Harry, et al.
Published: (2024)
by: Dong, Harry, et al.
Published: (2024)
LinkQ: An LLM-Assisted Visual Interface for Knowledge Graph Question-Answering
by: Li, Harry, et al.
Published: (2024)
by: Li, Harry, et al.
Published: (2024)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
by: Zhu, Xudong, et al.
Published: (2025)
by: Zhu, Xudong, et al.
Published: (2025)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
by: Wang, Shouren, et al.
Published: (2025)
by: Wang, Shouren, et al.
Published: (2025)
IPAD: Inverse Prompt for AI Detection - A Robust and Interpretable LLM-Generated Text Detector
by: Chen, Zheng, et al.
Published: (2025)
by: Chen, Zheng, et al.
Published: (2025)
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
by: Nikdan, Mahdi, et al.
Published: (2024)
by: Nikdan, Mahdi, et al.
Published: (2024)
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)
by: Nikdan, Mahdi, et al.
Published: (2026)
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
by: Manvi, Rohin, et al.
Published: (2024)
by: Manvi, Rohin, et al.
Published: (2024)
Fine-Tuning Small Language Models (SLMs) for Autonomous Web-based Geographical Information Systems (AWebGIS)
by: Ashani, Mahdi Nazari, et al.
Published: (2025)
by: Ashani, Mahdi Nazari, et al.
Published: (2025)
When Can Transformers Count to n?
by: Yehudai, Gilad, et al.
Published: (2024)
by: Yehudai, Gilad, et al.
Published: (2024)
Can LLMs Follow Simple Rules?
by: Mu, Norman, et al.
Published: (2023)
by: Mu, Norman, et al.
Published: (2023)
Post-Trained MoE Can Skip Half Experts via Self-Distillation
by: Lv, Xingtai, et al.
Published: (2026)
by: Lv, Xingtai, et al.
Published: (2026)
Similar Items
-
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025) -
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
by: Yang, Yushi, et al.
Published: (2024) -
Negation Neglect: When models fail to learn negations in training
by: Mayne, Harry, et al.
Published: (2026) -
Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
by: Sondej, Filip, et al.
Published: (2025) -
Multi-view autoencoders for Fake News Detection
by: Pereira, Ingryd V. S. T., et al.
Published: (2025)