Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Balagansky, Nikita, Aksenov, Yaroslav, Laptev, Daniil, Kurochkin, Vadim, Gerasimov, Gleb, Koryagin, Nikita, Gavrilov, Daniil
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916780232933376
author Balagansky, Nikita
Aksenov, Yaroslav
Laptev, Daniil
Kurochkin, Vadim
Gerasimov, Gleb
Koryagin, Nikita
Gavrilov, Daniil
author_facet Balagansky, Nikita
Aksenov, Yaroslav
Laptev, Daniil
Kurochkin, Vadim
Gerasimov, Gleb
Koryagin, Nikita
Gavrilov, Daniil
contents Sparse Autoencoders (SAEs) have proven to be powerful tools for interpreting neural networks by decomposing hidden representations into disentangled, interpretable features via sparsity constraints. However, conventional SAEs are constrained by the fixed sparsity level chosen during training; meeting different sparsity requirements therefore demands separate models and increases the computational footprint during both training and evaluation. We introduce a novel training objective, \emph{HierarchicalTopK}, which trains a single SAE to optimise reconstructions across multiple sparsity levels simultaneously. Experiments with Gemma-2 2B demonstrate that our approach achieves Pareto-optimal trade-offs between sparsity and explained variance, outperforming traditional SAEs trained at individual sparsity levels. Further analysis shows that HierarchicalTopK preserves high interpretability scores even at higher sparsity. The proposed objective thus closes an important gap between flexibility and interpretability in SAE design.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24473
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
Balagansky, Nikita
Aksenov, Yaroslav
Laptev, Daniil
Kurochkin, Vadim
Gerasimov, Gleb
Koryagin, Nikita
Gavrilov, Daniil
Machine Learning
Artificial Intelligence
Sparse Autoencoders (SAEs) have proven to be powerful tools for interpreting neural networks by decomposing hidden representations into disentangled, interpretable features via sparsity constraints. However, conventional SAEs are constrained by the fixed sparsity level chosen during training; meeting different sparsity requirements therefore demands separate models and increases the computational footprint during both training and evaluation. We introduce a novel training objective, \emph{HierarchicalTopK}, which trains a single SAE to optimise reconstructions across multiple sparsity levels simultaneously. Experiments with Gemma-2 2B demonstrate that our approach achieves Pareto-optimal trade-offs between sparsity and explained variance, outperforming traditional SAEs trained at individual sparsity levels. Further analysis shows that HierarchicalTopK preserves high interpretability scores even at higher sparsity. The proposed objective thus closes an important gap between flexibility and interpretability in SAE design.
title Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.24473