Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
Fuente:
arXiv
Saved in:
| Main Authors: | Costa, Valérie, Fel, Thomas, Lubana, Ekdeep Singh, Tolooshams, Bahareh, Ba, Demba |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
by: Costa, Valérie, et al.
Published: (2025)
by: Costa, Valérie, et al.
Published: (2025)
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
by: Hindupur, Sai Sumedh R., et al.
Published: (2025)
by: Hindupur, Sai Sumedh R., et al.
Published: (2025)
Discriminative reconstruction via simultaneous dense and sparse coding
by: Tasissa, Abiy, et al.
Published: (2020)
by: Tasissa, Abiy, et al.
Published: (2020)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
by: Tolooshams, Bahareh, et al.
Published: (2025)
by: Tolooshams, Bahareh, et al.
Published: (2025)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
by: Bohacek, Matyas, et al.
Published: (2025)
by: Bohacek, Matyas, et al.
Published: (2025)
Do Sparse Autoencoders Capture Concept Manifolds?
by: Bhalla, Usha, et al.
Published: (2026)
by: Bhalla, Usha, et al.
Published: (2026)
Abrupt Learning in Transformers: A Case Study on Matrix Completion
by: Gopalani, Pulkit, et al.
Published: (2024)
by: Gopalani, Pulkit, et al.
Published: (2024)
Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
by: Grosso, Gaia, et al.
Published: (2025)
by: Grosso, Gaia, et al.
Published: (2025)
Priors in Time: Missing Inductive Biases for Language Model Interpretability
by: Lubana, Ekdeep Singh, et al.
Published: (2025)
by: Lubana, Ekdeep Singh, et al.
Published: (2025)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
by: Jaipersaud, Brandon, et al.
Published: (2025)
by: Jaipersaud, Brandon, et al.
Published: (2025)
Analyzing (In)Abilities of SAEs via Formal Languages
by: Menon, Abhinav, et al.
Published: (2024)
by: Menon, Abhinav, et al.
Published: (2024)
Block-Recurrent Dynamics in Vision Transformers
by: Jacobs, Mozes, et al.
Published: (2025)
by: Jacobs, Mozes, et al.
Published: (2025)
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
by: Okawa, Maya, et al.
Published: (2023)
by: Okawa, Maya, et al.
Published: (2023)
Diffusion State-Guided Projected Gradient for Inverse Problems
by: Zirvi, Rayhan, et al.
Published: (2024)
by: Zirvi, Rayhan, et al.
Published: (2024)
A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language
by: Lubana, Ekdeep Singh, et al.
Published: (2024)
by: Lubana, Ekdeep Singh, et al.
Published: (2024)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks
by: Ramesh, Rahul, et al.
Published: (2023)
by: Ramesh, Rahul, et al.
Published: (2023)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
by: Thasarathan, Harrish, et al.
Published: (2025)
by: Thasarathan, Harrish, et al.
Published: (2025)
Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing
by: Nishi, Kento, et al.
Published: (2024)
by: Nishi, Kento, et al.
Published: (2024)
Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
Clustering Inductive Biases with Unrolled Networks
by: Huml, Jonathan, et al.
Published: (2023)
by: Huml, Jonathan, et al.
Published: (2023)
Swing-by Dynamics in Concept Learning and Compositional Generalization
by: Yang, Yongyi, et al.
Published: (2024)
by: Yang, Yongyi, et al.
Published: (2024)
Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models
by: Fel, Thomas, et al.
Published: (2025)
by: Fel, Thomas, et al.
Published: (2025)
The Impact of Off-Policy Training Data on Probe Generalisation
by: Kirch, Nathalie, et al.
Published: (2025)
by: Kirch, Nathalie, et al.
Published: (2025)
Weighed l1 on the simplex: Compressive sensing meets locality
by: Tasissa, Abiy, et al.
Published: (2021)
by: Tasissa, Abiy, et al.
Published: (2021)
Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability
by: Prasad, Aaditya Vikram, et al.
Published: (2026)
by: Prasad, Aaditya Vikram, et al.
Published: (2026)
Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
by: Fel, Thomas, et al.
Published: (2025)
by: Fel, Thomas, et al.
Published: (2025)
Detecting High-Stakes Interactions with Activation Probes
by: McKenzie, Alex, et al.
Published: (2025)
by: McKenzie, Alex, et al.
Published: (2025)
Quantum Sparse Recovery and Quantum Orthogonal Matching Pursuit
by: Bellante, Armando, et al.
Published: (2025)
by: Bellante, Armando, et al.
Published: (2025)
In-Context Learning Strategies Emerge Rationally
by: Wurgaft, Daniel, et al.
Published: (2025)
by: Wurgaft, Daniel, et al.
Published: (2025)
In-Context Learning Dynamics with Random Binary Sequences
by: Bigelow, Eric J., et al.
Published: (2023)
by: Bigelow, Eric J., et al.
Published: (2023)
Evaluating and Designing Sparse Autoencoders by Approximating Quasi-Orthogonality
by: Lee, Sewoong, et al.
Published: (2025)
by: Lee, Sewoong, et al.
Published: (2025)
Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
by: Wurgaft, Daniel, et al.
Published: (2026)
by: Wurgaft, Daniel, et al.
Published: (2026)
Fourier Neural Operators for Learning Dynamics in Quantum Spin Systems
by: Shah, Freya, et al.
Published: (2024)
by: Shah, Freya, et al.
Published: (2024)
Origins of Creativity in Attention-Based Diffusion Models
by: Finn, Emma, et al.
Published: (2025)
by: Finn, Emma, et al.
Published: (2025)
Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation Model
by: Khona, Mikail, et al.
Published: (2024)
by: Khona, Mikail, et al.
Published: (2024)
What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
by: Jain, Samyak, et al.
Published: (2024)
by: Jain, Samyak, et al.
Published: (2024)
Evaluating Sparse Autoencoders for Monosemantic Representation
by: Fereidouni, Moghis, et al.
Published: (2025)
by: Fereidouni, Moghis, et al.
Published: (2025)
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
by: Jain, Samyak, et al.
Published: (2023)
by: Jain, Samyak, et al.
Published: (2023)
Similar Items
-
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
by: Costa, Valérie, et al.
Published: (2025) -
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
by: Hindupur, Sai Sumedh R., et al.
Published: (2025) -
Discriminative reconstruction via simultaneous dense and sparse coding
by: Tasissa, Abiy, et al.
Published: (2020) -
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
by: Tolooshams, Bahareh, et al.
Published: (2025) -
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
by: Bohacek, Matyas, et al.
Published: (2025)