Concept-SAE: Active Causal Probing of Visual Model Behavior

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Jianrong, Chen, Muxi, Zhao, Chenchen, Xu, Qiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914058038411264
author Ding, Jianrong
Chen, Muxi
Zhao, Chenchen
Xu, Qiang
author_facet Ding, Jianrong
Chen, Muxi
Zhao, Chenchen
Xu, Qiang
contents Standard Sparse Autoencoders (SAEs) excel at discovering a dictionary of a model's learned features, offering a powerful observational lens. However, the ambiguous and ungrounded nature of these features makes them unreliable instruments for the active, causal probing of model behavior. To solve this, we introduce Concept-SAE, a framework that forges semantically grounded concept tokens through a novel hybrid disentanglement strategy. We first quantitatively demonstrate that our dual-supervision approach produces tokens that are remarkably faithful and spatially localized, outperforming alternative methods in disentanglement. This validated fidelity enables two critical applications: (1) we probe the causal link between internal concepts and predictions via direct intervention, and (2) we probe the model's failure modes by systematically localizing adversarial vulnerabilities to specific layers. Concept-SAE provides a validated blueprint for moving beyond correlational interpretation to the mechanistic, causal probing of model behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22015
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Concept-SAE: Active Causal Probing of Visual Model Behavior
Ding, Jianrong
Chen, Muxi
Zhao, Chenchen
Xu, Qiang
Machine Learning
Standard Sparse Autoencoders (SAEs) excel at discovering a dictionary of a model's learned features, offering a powerful observational lens. However, the ambiguous and ungrounded nature of these features makes them unreliable instruments for the active, causal probing of model behavior. To solve this, we introduce Concept-SAE, a framework that forges semantically grounded concept tokens through a novel hybrid disentanglement strategy. We first quantitatively demonstrate that our dual-supervision approach produces tokens that are remarkably faithful and spatially localized, outperforming alternative methods in disentanglement. This validated fidelity enables two critical applications: (1) we probe the causal link between internal concepts and predictions via direct intervention, and (2) we probe the model's failure modes by systematically localizing adversarial vulnerabilities to specific layers. Concept-SAE provides a validated blueprint for moving beyond correlational interpretation to the mechanistic, causal probing of model behavior.
title Concept-SAE: Active Causal Probing of Visual Model Behavior
topic Machine Learning
url https://arxiv.org/abs/2509.22015