Interpretable and Steerable Concept Bottleneck Sparse Autoencoders

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kulkarni, Akshay, Weng, Tsui-Wei, Narayanaswamy, Vivek, Liu, Shusen, Sakla, Wesam A., Thopalli, Kowshik
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915900623421440
author Kulkarni, Akshay
Weng, Tsui-Wei
Narayanaswamy, Vivek
Liu, Shusen
Sakla, Wesam A.
Thopalli, Kowshik
author_facet Kulkarni, Akshay
Weng, Tsui-Wei
Narayanaswamy, Vivek
Liu, Shusen
Sakla, Wesam A.
Thopalli, Kowshik
contents Sparse autoencoders (SAEs) promise a unified approach for mechanistic interpretability, concept discovery, and model steering in LLMs and LVLMs. However, realizing this potential requires learned features to be both interpretable and steerable. To that end, we introduce two new computationally inexpensive interpretability and steerability metrics for a systematic analysis of LVLM SAEs. This uncovers two observations; (i) a majority of SAE neurons exhibit either low interpretability or low steerability or both, rendering them ineffective for downstream use; and (ii) user-desired concepts are often absent in the SAE, thus limiting their practical utility. To address these limitations, we propose Concept Bottleneck Sparse Autoencoders (CB-SAE) - a novel post-hoc framework that prunes low-utility neurons and augments the latent space with a lightweight concept bottleneck aligned to a user-defined concept set. The resulting CB-SAE improves interpretability by +32.1% and steerability by +14.5% across LVLMs and image generation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10805
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
Kulkarni, Akshay
Weng, Tsui-Wei
Narayanaswamy, Vivek
Liu, Shusen
Sakla, Wesam A.
Thopalli, Kowshik
Machine Learning
Computer Vision and Pattern Recognition
Sparse autoencoders (SAEs) promise a unified approach for mechanistic interpretability, concept discovery, and model steering in LLMs and LVLMs. However, realizing this potential requires learned features to be both interpretable and steerable. To that end, we introduce two new computationally inexpensive interpretability and steerability metrics for a systematic analysis of LVLM SAEs. This uncovers two observations; (i) a majority of SAE neurons exhibit either low interpretability or low steerability or both, rendering them ineffective for downstream use; and (ii) user-desired concepts are often absent in the SAE, thus limiting their practical utility. To address these limitations, we propose Concept Bottleneck Sparse Autoencoders (CB-SAE) - a novel post-hoc framework that prunes low-utility neurons and augments the latent space with a lightweight concept bottleneck aligned to a user-defined concept set. The resulting CB-SAE improves interpretability by +32.1% and steerability by +14.5% across LVLMs and image generation tasks.
title Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.10805