Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lim, Sungjun, Kim, Heedong, Lee, Andrew, Song, Kyungwoo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917518285733888
author Lim, Sungjun
Kim, Heedong
Lee, Andrew
Song, Kyungwoo
author_facet Lim, Sungjun
Kim, Heedong
Lee, Andrew
Song, Kyungwoo
contents Mechanistic interpretability aims to explain a model's behavior by identifying causally responsible internal structures. Dictionary-based explainers such as sparse autoencoders and transcoders are a primary tool, but their faithfulness under out-of-distribution (OOD) shift has received little systematic attention. We show that distribution shift rotates the subspace that the model actively uses, misaligning the explainer's dictionary trained on in-distribution (ID) activations. We formalize this misalignment as the faithfulness gap, a geometric distance between the ID dictionary and the OOD-active subspace, and show that it controls OOD faithfulness degradation. To reduce this gap, we propose the Geometry-Adaptive Explainer (GAE), which realigns the explainer's dictionary with the OOD-active subspace while preserving the original feature structure. This requires only unlabeled OOD activations and no gradient updates. We prove that GAE improves over the unadapted ID explainer, with excess loss bounded quadratically by the second-moment shift. Empirically, GAE even matches or surpasses all training-based baselines in causal faithfulness across multiple models and OOD settings.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21849
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift
Lim, Sungjun
Kim, Heedong
Lee, Andrew
Song, Kyungwoo
Machine Learning
Computation and Language
Mechanistic interpretability aims to explain a model's behavior by identifying causally responsible internal structures. Dictionary-based explainers such as sparse autoencoders and transcoders are a primary tool, but their faithfulness under out-of-distribution (OOD) shift has received little systematic attention. We show that distribution shift rotates the subspace that the model actively uses, misaligning the explainer's dictionary trained on in-distribution (ID) activations. We formalize this misalignment as the faithfulness gap, a geometric distance between the ID dictionary and the OOD-active subspace, and show that it controls OOD faithfulness degradation. To reduce this gap, we propose the Geometry-Adaptive Explainer (GAE), which realigns the explainer's dictionary with the OOD-active subspace while preserving the original feature structure. This requires only unlabeled OOD activations and no gradient updates. We prove that GAE improves over the unadapted ID explainer, with excess loss bounded quadratically by the second-moment shift. Empirically, GAE even matches or surpasses all training-based baselines in causal faithfulness across multiple models and OOD settings.
title Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2605.21849