Sparse Autoencoder Features for Classifications and Transferability

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gallifant, Jack, Chen, Shan, Sasse, Kuleen, Aerts, Hugo, Hartvigsen, Thomas, Bitterman, Danielle S.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910007952408576
author Gallifant, Jack
Chen, Shan
Sasse, Kuleen
Aerts, Hugo
Hartvigsen, Thomas
Bitterman, Danielle S.
author_facet Gallifant, Jack
Chen, Shan
Sasse, Kuleen
Aerts, Hugo
Hartvigsen, Thomas
Bitterman, Danielle S.
contents Sparse Autoencoders (SAEs) provide potentials for uncovering structured, human-interpretable representations in Large Language Models (LLMs), making them a crucial tool for transparent and controllable AI systems. We systematically analyze SAE for interpretable feature extraction from LLMs in safety-critical classification tasks. Our framework evaluates (1) model-layer selection and scaling properties, (2) SAE architectural configurations, including width and pooling strategies, and (3) the effect of binarizing continuous SAE activations. SAE-derived features achieve macro F1 > 0.8, outperforming hidden-state and BoW baselines while demonstrating cross-model transfer from Gemma 2 2B to 9B-IT models. These features generalize in a zero-shot manner to cross-lingual toxicity detection and visual classification tasks. Our analysis highlights the significant impact of pooling strategies and binarization thresholds, showing that binarization offers an efficient alternative to traditional feature selection while maintaining or improving performance. These findings establish new best practices for SAE-based interpretability and enable scalable, transparent deployment of LLMs in real-world applications. Full repo: https://github.com/shan23chen/MOSAIC.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11367
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sparse Autoencoder Features for Classifications and Transferability
Gallifant, Jack
Chen, Shan
Sasse, Kuleen
Aerts, Hugo
Hartvigsen, Thomas
Bitterman, Danielle S.
Machine Learning
Artificial Intelligence
Computation and Language
Sparse Autoencoders (SAEs) provide potentials for uncovering structured, human-interpretable representations in Large Language Models (LLMs), making them a crucial tool for transparent and controllable AI systems. We systematically analyze SAE for interpretable feature extraction from LLMs in safety-critical classification tasks. Our framework evaluates (1) model-layer selection and scaling properties, (2) SAE architectural configurations, including width and pooling strategies, and (3) the effect of binarizing continuous SAE activations. SAE-derived features achieve macro F1 > 0.8, outperforming hidden-state and BoW baselines while demonstrating cross-model transfer from Gemma 2 2B to 9B-IT models. These features generalize in a zero-shot manner to cross-lingual toxicity detection and visual classification tasks. Our analysis highlights the significant impact of pooling strategies and binarization thresholds, showing that binarization offers an efficient alternative to traditional feature selection while maintaining or improving performance. These findings establish new best practices for SAE-based interpretability and enable scalable, transparent deployment of LLMs in real-world applications. Full repo: https://github.com/shan23chen/MOSAIC.
title Sparse Autoencoder Features for Classifications and Transferability
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.11367