Do Sparse Autoencoders Generalize? A Case Study of Answerability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Heindrich, Lovis, Torr, Philip, Barez, Fazl, Thost, Veronika
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911139254763520
author Heindrich, Lovis
Torr, Philip
Barez, Fazl
Thost, Veronika
author_facet Heindrich, Lovis
Torr, Philip
Barez, Fazl
Thost, Veronika
contents Sparse autoencoders (SAEs) have emerged as a promising approach in language model interpretability, offering unsupervised extraction of sparse features. For interpretability methods to succeed, they must identify abstract features across domains, and these features can often manifest differently in each context. We examine this through "answerability" - a model's ability to recognize answerable questions. We extensively evaluate SAE feature generalization across diverse, partly self-constructed answerability datasets for Gemma 2 SAEs. Our analysis reveals that residual stream probes outperform SAE features within domains, but generalization performance differs sharply. SAE features show inconsistent out-of-domain transfer, with performance varying from almost random to outperforming residual stream probes. Overall, this demonstrates the need for robust evaluation methods and quantitative approaches to predict feature generalization in SAE-based interpretability.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19964
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do Sparse Autoencoders Generalize? A Case Study of Answerability
Heindrich, Lovis
Torr, Philip
Barez, Fazl
Thost, Veronika
Machine Learning
Sparse autoencoders (SAEs) have emerged as a promising approach in language model interpretability, offering unsupervised extraction of sparse features. For interpretability methods to succeed, they must identify abstract features across domains, and these features can often manifest differently in each context. We examine this through "answerability" - a model's ability to recognize answerable questions. We extensively evaluate SAE feature generalization across diverse, partly self-constructed answerability datasets for Gemma 2 SAEs. Our analysis reveals that residual stream probes outperform SAE features within domains, but generalization performance differs sharply. SAE features show inconsistent out-of-domain transfer, with performance varying from almost random to outperforming residual stream probes. Overall, this demonstrates the need for robust evaluation methods and quantitative approaches to predict feature generalization in SAE-based interpretability.
title Do Sparse Autoencoders Generalize? A Case Study of Answerability
topic Machine Learning
url https://arxiv.org/abs/2502.19964