Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
Fuente:
arXiv
Saved in:
| Main Authors: | Kantamneni, Subhash, Engels, Joshua, Rajamanoharan, Senthooran, Tegmark, Max, Nanda, Neel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Language Models Use Trigonometry to Do Addition
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
Improving Dictionary Learning with Gated Sparse Autoencoders
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
Scaling Laws For Scalable Oversight
by: Engels, Joshua, et al.
Published: (2025)
by: Engels, Joshua, et al.
Published: (2025)
Convergent Linear Representations of Emergent Misalignment
by: Soligo, Anna, et al.
Published: (2025)
by: Soligo, Anna, et al.
Published: (2025)
How Do Transformers "Do" Physics? Investigating the Simple Harmonic Oscillator
by: Kantamneni, Subhash, et al.
Published: (2024)
by: Kantamneni, Subhash, et al.
Published: (2024)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
by: Ferrando, Javier, et al.
Published: (2024)
by: Ferrando, Javier, et al.
Published: (2024)
Dense SAE Latents Are Features, Not Bugs
by: Sun, Xiaoqing, et al.
Published: (2025)
by: Sun, Xiaoqing, et al.
Published: (2025)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
by: Lieberum, Tom, et al.
Published: (2024)
by: Lieberum, Tom, et al.
Published: (2024)
The Geometry of Concepts: Sparse Autoencoder Feature Structure
by: Li, Yuxiao, et al.
Published: (2024)
by: Li, Yuxiao, et al.
Published: (2024)
Model Organisms for Emergent Misalignment
by: Turner, Edward, et al.
Published: (2025)
by: Turner, Edward, et al.
Published: (2025)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
by: Macar, Uzay, et al.
Published: (2025)
by: Macar, Uzay, et al.
Published: (2025)
Decomposing The Dark Matter of Sparse Autoencoders
by: Engels, Joshua, et al.
Published: (2024)
by: Engels, Joshua, et al.
Published: (2024)
BatchTopK Sparse Autoencoders
by: Bussmann, Bart, et al.
Published: (2024)
by: Bussmann, Bart, et al.
Published: (2024)
Low-Rank Adapting Models for Sparse Autoencoders
by: Chen, Matthew, et al.
Published: (2025)
by: Chen, Matthew, et al.
Published: (2025)
How Well Do Models Follow Their Constitutions?
by: Jakkli, Arya, et al.
Published: (2026)
by: Jakkli, Arya, et al.
Published: (2026)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
by: Bussmann, Bart, et al.
Published: (2025)
by: Bussmann, Bart, et al.
Published: (2025)
Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
by: Casademunt, Helena, et al.
Published: (2025)
by: Casademunt, Helena, et al.
Published: (2025)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
by: Arcuschin, Iván, et al.
Published: (2025)
by: Arcuschin, Iván, et al.
Published: (2025)
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
by: Jiang, Nick, et al.
Published: (2025)
by: Jiang, Nick, et al.
Published: (2025)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
by: Wang, Atticus, et al.
Published: (2025)
by: Wang, Atticus, et al.
Published: (2025)
OptPDE: Discovering Novel Integrable Systems via AI-Human Collaboration
by: Kantamneni, Subhash, et al.
Published: (2024)
by: Kantamneni, Subhash, et al.
Published: (2024)
Sparse Autoencoders Do Not Find Canonical Units of Analysis
by: Leask, Patrick, et al.
Published: (2025)
by: Leask, Patrick, et al.
Published: (2025)
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
by: Cywiński, Bartosz, et al.
Published: (2025)
by: Cywiński, Bartosz, et al.
Published: (2025)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
by: Soligo, Anna, et al.
Published: (2026)
by: Soligo, Anna, et al.
Published: (2026)
Efficient Dictionary Learning with Switch Sparse Autoencoders
by: Mudide, Anish, et al.
Published: (2024)
by: Mudide, Anish, et al.
Published: (2024)
Are Sparse Autoencoders Useful for Java Function Bug Detection?
by: Melo, Rui, et al.
Published: (2025)
by: Melo, Rui, et al.
Published: (2025)
Building Production-Ready Probes For Gemini
by: Kramár, János, et al.
Published: (2026)
by: Kramár, János, et al.
Published: (2026)
Subliminal Learning Is Steering Vector Distillation
by: Blank, Camila, et al.
Published: (2026)
by: Blank, Camila, et al.
Published: (2026)
ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions
by: Poduval, Prathyush, et al.
Published: (2026)
by: Poduval, Prathyush, et al.
Published: (2026)
Investigating Representation Universality: Case Study on Genealogical Representations
by: Baek, David D., et al.
Published: (2024)
by: Baek, David D., et al.
Published: (2024)
Sparse Autoencoders, Again?
by: Lu, Yin, et al.
Published: (2025)
by: Lu, Yin, et al.
Published: (2025)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
by: Makelov, Aleksandar, et al.
Published: (2024)
by: Makelov, Aleksandar, et al.
Published: (2024)
Are Sparse Autoencoder Benchmarks Reliable?
by: Chanin, David
Published: (2026)
by: Chanin, David
Published: (2026)
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
by: Ayonrinde, Kola
Published: (2024)
by: Ayonrinde, Kola
Published: (2024)
Improving Sparse Autoencoder with Dynamic Attention
by: Wang, Dongsheng, et al.
Published: (2026)
by: Wang, Dongsheng, et al.
Published: (2026)
SparseJEPA: Sparse Representation Learning of Joint Embedding Predictive Architectures
by: Hartman, Max, et al.
Published: (2025)
by: Hartman, Max, et al.
Published: (2025)
Empirical Evaluation of Progressive Coding for Sparse Autoencoders
by: Peter, Hans, et al.
Published: (2025)
by: Peter, Hans, et al.
Published: (2025)
On the transferability of Sparse Autoencoders for interpreting compressed models
by: Gupte, Suchit, et al.
Published: (2025)
by: Gupte, Suchit, et al.
Published: (2025)
Data Whitening Improves Sparse Autoencoder Learning
by: Saraswatula, Ashwin, et al.
Published: (2025)
by: Saraswatula, Ashwin, et al.
Published: (2025)
Similar Items
-
Language Models Use Trigonometry to Do Addition
by: Kantamneni, Subhash, et al.
Published: (2025) -
Improving Dictionary Learning with Gated Sparse Autoencoders
by: Rajamanoharan, Senthooran, et al.
Published: (2024) -
Scaling Laws For Scalable Oversight
by: Engels, Joshua, et al.
Published: (2025) -
Convergent Linear Representations of Emergent Misalignment
by: Soligo, Anna, et al.
Published: (2025) -
How Do Transformers "Do" Physics? Investigating the Simple Harmonic Oscillator
by: Kantamneni, Subhash, et al.
Published: (2024)