Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Costa, Valérie, Fel, Thomas, Lubana, Ekdeep Singh, Tolooshams, Bahareh, Ba, Demba
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915594250485760
author Costa, Valérie
Fel, Thomas
Lubana, Ekdeep Singh
Tolooshams, Bahareh
Ba, Demba
author_facet Costa, Valérie
Fel, Thomas
Lubana, Ekdeep Singh
Tolooshams, Bahareh
Ba, Demba
contents Sparse autoencoders (SAEs) have recently become central tools for interpretability, leveraging dictionary learning principles to extract sparse, interpretable features from neural representations whose underlying structure is typically unknown. This paper evaluates SAEs in a controlled setting using MNIST, which reveals that current shallow architectures implicitly rely on a quasi-orthogonality assumption that limits the ability to extract correlated features. To move beyond this, we compare them with an iterative SAE that unrolls Matching Pursuit (MP-SAE), enabling the residual-guided extraction of correlated features that arise in hierarchical settings such as handwritten digit generation while guaranteeing monotonic improvement of the reconstruction as more atoms are selected.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05239
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
Costa, Valérie
Fel, Thomas
Lubana, Ekdeep Singh
Tolooshams, Bahareh
Ba, Demba
Machine Learning
Sparse autoencoders (SAEs) have recently become central tools for interpretability, leveraging dictionary learning principles to extract sparse, interpretable features from neural representations whose underlying structure is typically unknown. This paper evaluates SAEs in a controlled setting using MNIST, which reveals that current shallow architectures implicitly rely on a quasi-orthogonality assumption that limits the ability to extract correlated features. To move beyond this, we compare them with an iterative SAE that unrolls Matching Pursuit (MP-SAE), enabling the residual-guided extraction of correlated features that arise in hierarchical settings such as handwritten digit generation while guaranteeing monotonic improvement of the reconstruction as more atoms are selected.
title Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
topic Machine Learning
url https://arxiv.org/abs/2506.05239