Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ayonrinde, Kola, Pearce, Michael T., Sharkey, Lee
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909349836750848
author Ayonrinde, Kola
Pearce, Michael T.
Sharkey, Lee
author_facet Ayonrinde, Kola
Pearce, Michael T.
Sharkey, Lee
contents Sparse Autoencoders (SAEs) have emerged as a useful tool for interpreting the internal representations of neural networks. However, naively optimising SAEs for reconstruction loss and sparsity results in a preference for SAEs that are extremely wide and sparse. We present an information-theoretic framework for interpreting SAEs as lossy compression algorithms for communicating explanations of neural activations. We appeal to the Minimal Description Length (MDL) principle to motivate explanations of activations which are both accurate and concise. We further argue that interpretable SAEs require an additional property, "independent additivity": features should be able to be understood separately. We demonstrate an example of applying our MDL-inspired framework by training SAEs on MNIST handwritten digits and find that SAE features representing significant line segments are optimal, as opposed to SAEs with features for memorised digits from the dataset or small digit fragments. We argue that using MDL rather than sparsity may avoid potential pitfalls with naively maximising sparsity such as undesirable feature splitting and that this framework naturally suggests new hierarchical SAE architectures which provide more concise explanations.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11179
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
Ayonrinde, Kola
Pearce, Michael T.
Sharkey, Lee
Machine Learning
Artificial Intelligence
Information Theory
Sparse Autoencoders (SAEs) have emerged as a useful tool for interpreting the internal representations of neural networks. However, naively optimising SAEs for reconstruction loss and sparsity results in a preference for SAEs that are extremely wide and sparse. We present an information-theoretic framework for interpreting SAEs as lossy compression algorithms for communicating explanations of neural activations. We appeal to the Minimal Description Length (MDL) principle to motivate explanations of activations which are both accurate and concise. We further argue that interpretable SAEs require an additional property, "independent additivity": features should be able to be understood separately. We demonstrate an example of applying our MDL-inspired framework by training SAEs on MNIST handwritten digits and find that SAE features representing significant line segments are optimal, as opposed to SAEs with features for memorised digits from the dataset or small digit fragments. We argue that using MDL rather than sparsity may avoid potential pitfalls with naively maximising sparsity such as undesirable feature splitting and that this framework naturally suggests new hierarchical SAE architectures which provide more concise explanations.
title Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
topic Machine Learning
Artificial Intelligence
Information Theory
url https://arxiv.org/abs/2410.11179