Transcoders Beat Sparse Autoencoders for Interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Paulo, Gonçalo, Shabalin, Stepan, Belrose, Nora |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Binary Sparse Coding for Interpretability
by: Quirke, Lucia, et al.
Published: (2025)
by: Quirke, Lucia, et al.
Published: (2025)
Sparse Autoencoders Trained on the Same Data Learn Different Features
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Evaluating SAE interpretability without explanations
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Does Transformer Interpretability Transfer to RNNs?
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Partially Rewriting a Transformer in Natural Language
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Automatically Interpreting Millions of Features in Large Language Models
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Estimating the Probability of Sampling a Trained Neural Network at Random
by: Scherlis, Adam, et al.
Published: (2025)
by: Scherlis, Adam, et al.
Published: (2025)
Slowing Learning by Erasing Simple Features
by: Quirke, Lucia, et al.
Published: (2025)
by: Quirke, Lucia, et al.
Published: (2025)
Converting MLPs into Polynomials in Closed Form
by: Belrose, Nora, et al.
Published: (2025)
by: Belrose, Nora, et al.
Published: (2025)
Balancing Label Quantity and Quality for Scalable Elicitation
by: Mallen, Alex, et al.
Published: (2024)
by: Mallen, Alex, et al.
Published: (2024)
Understanding Gradient Descent through the Training Jacobian
by: Belrose, Nora, et al.
Published: (2024)
by: Belrose, Nora, et al.
Published: (2024)
Examining Two Hop Reasoning Through Information Content Scaling
by: Johnston, David, et al.
Published: (2025)
by: Johnston, David, et al.
Published: (2025)
Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning
by: Shabalin, Stepan, et al.
Published: (2025)
by: Shabalin, Stepan, et al.
Published: (2025)
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
by: Korznikov, Anton, et al.
Published: (2026)
by: Korznikov, Anton, et al.
Published: (2026)
Transcoders Find Interpretable LLM Feature Circuits
by: Dunefsky, Jacob, et al.
Published: (2024)
by: Dunefsky, Jacob, et al.
Published: (2024)
Refusal in LLMs is an Affine Function
by: Marshall, Thomas, et al.
Published: (2024)
by: Marshall, Thomas, et al.
Published: (2024)
Transcoder-based Circuit Analysis for Interpretable Single-Cell Foundation Models
by: Hosokawa, Sosuke, et al.
Published: (2025)
by: Hosokawa, Sosuke, et al.
Published: (2025)
Interpretable Reward Model via Sparse Autoencoder
by: Zhang, Shuyi, et al.
Published: (2025)
by: Zhang, Shuyi, et al.
Published: (2025)
Interpreting Attention Layer Outputs with Sparse Autoencoders
by: Kissane, Connor, et al.
Published: (2024)
by: Kissane, Connor, et al.
Published: (2024)
Mechanistic Anomaly Detection for "Quirky" Language Models
by: Johnston, David O., et al.
Published: (2025)
by: Johnston, David O., et al.
Published: (2025)
Route Sparse Autoencoder to Interpret Large Language Models
by: Shi, Wei, et al.
Published: (2025)
by: Shi, Wei, et al.
Published: (2025)
Step-Level Sparse Autoencoder for Reasoning Process Interpretation
by: Yang, Xuan, et al.
Published: (2026)
by: Yang, Xuan, et al.
Published: (2026)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
by: Makelov, Aleksandar, et al.
Published: (2024)
by: Makelov, Aleksandar, et al.
Published: (2024)
Beyond Dense States: Elevating Sparse Transcoders to Active Operators for Latent Reasoning
by: Wang, Yadong, et al.
Published: (2026)
by: Wang, Yadong, et al.
Published: (2026)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
by: Erdogan, Ege, et al.
Published: (2025)
by: Erdogan, Ege, et al.
Published: (2025)
Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders
by: Ye, Mengyu, et al.
Published: (2025)
by: Ye, Mengyu, et al.
Published: (2025)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
by: Marks, Luke, et al.
Published: (2024)
by: Marks, Luke, et al.
Published: (2024)
Interpretable Company Similarity with Sparse Autoencoders
by: Molinari, Marco, et al.
Published: (2024)
by: Molinari, Marco, et al.
Published: (2024)
Neural Networks Learn Statistics of Increasing Complexity
by: Belrose, Nora, et al.
Published: (2024)
by: Belrose, Nora, et al.
Published: (2024)
Scaling sparse feature circuit finding for in-context learning
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
by: Kurochkin, Vadim, et al.
Published: (2025)
by: Kurochkin, Vadim, et al.
Published: (2025)
Eliciting Latent Knowledge from Quirky Language Models
by: Mallen, Alex, et al.
Published: (2023)
by: Mallen, Alex, et al.
Published: (2023)
Transcoder Adapters for Reasoning-Model Diffing
by: Hu, Nathan, et al.
Published: (2026)
by: Hu, Nathan, et al.
Published: (2026)
Interpreting CLIP with Hierarchical Sparse Autoencoders
by: Zaigrajew, Vladimir, et al.
Published: (2025)
by: Zaigrajew, Vladimir, et al.
Published: (2025)
Interpreting and Steering Protein Language Models through Sparse Autoencoders
by: Garcia, Edith Natalia Villegas, et al.
Published: (2025)
by: Garcia, Edith Natalia Villegas, et al.
Published: (2025)
Interpreting CFD Surrogates through Sparse Autoencoders
by: Hu, Yeping, et al.
Published: (2025)
by: Hu, Yeping, et al.
Published: (2025)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
by: Tolooshams, Bahareh, et al.
Published: (2025)
by: Tolooshams, Bahareh, et al.
Published: (2025)
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
by: Kulkarni, Akshay, et al.
Published: (2025)
by: Kulkarni, Akshay, et al.
Published: (2025)
Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
by: O'Neill, Charles, et al.
Published: (2025)
by: O'Neill, Charles, et al.
Published: (2025)
XNNTab -- Interpretable Neural Networks for Tabular Data using Sparse Autoencoders
by: Elhadri, Khawla, et al.
Published: (2025)
by: Elhadri, Khawla, et al.
Published: (2025)
Similar Items
-
Binary Sparse Coding for Interpretability
by: Quirke, Lucia, et al.
Published: (2025) -
Sparse Autoencoders Trained on the Same Data Learn Different Features
by: Paulo, Gonçalo, et al.
Published: (2025) -
Evaluating SAE interpretability without explanations
by: Paulo, Gonçalo, et al.
Published: (2025) -
Does Transformer Interpretability Transfer to RNNs?
by: Paulo, Gonçalo, et al.
Published: (2024) -
Partially Rewriting a Transformer in Natural Language
by: Paulo, Gonçalo, et al.
Published: (2025)