Toward Identifiable Sparse Autoencoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nelson, Walter, Karaletsos, Theofanis, Locatello, Francesco
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910272078217216
author Nelson, Walter
Karaletsos, Theofanis
Locatello, Francesco
author_facet Nelson, Walter
Karaletsos, Theofanis
Locatello, Francesco
contents Recently, sparse autoencoders (SAEs) have emerged as an attractive tool for interpreting and interacting with representations in practical neural networks. While it is common empirical folklore, we also show theoretically that SAEs are highly unstable: different training runs are likely to produce different concept dictionaries and sparse codes. We characterize the model properties that hinder the stability of real-world SAEs, and address each of these problems through minimal changes to the architecture and training procedure. Together, these changes yield two versions of an \textbf{i}dentifiable SAE (iSAE), a variant of the standard TopK SAE with lower reconstruction error and improved stability. We explain this improvement theoretically by connecting SAEs with traditional dictionary learning approaches, and show that the dictionaries learned in practice satisfy an approximate restricted isometry condition, rendering the corresponding sparse codes in those models near-identifiable.
format Preprint
id arxiv_https___arxiv_org_abs_2605_31245
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Toward Identifiable Sparse Autoencoders
Nelson, Walter
Karaletsos, Theofanis
Locatello, Francesco
Machine Learning
Recently, sparse autoencoders (SAEs) have emerged as an attractive tool for interpreting and interacting with representations in practical neural networks. While it is common empirical folklore, we also show theoretically that SAEs are highly unstable: different training runs are likely to produce different concept dictionaries and sparse codes. We characterize the model properties that hinder the stability of real-world SAEs, and address each of these problems through minimal changes to the architecture and training procedure. Together, these changes yield two versions of an \textbf{i}dentifiable SAE (iSAE), a variant of the standard TopK SAE with lower reconstruction error and improved stability. We explain this improvement theoretically by connecting SAEs with traditional dictionary learning approaches, and show that the dictionaries learned in practice satisfy an approximate restricted isometry condition, rendering the corresponding sparse codes in those models near-identifiable.
title Toward Identifiable Sparse Autoencoders
topic Machine Learning
url https://arxiv.org/abs/2605.31245