Zero-Shot Textual Explanations via Translating Decision-Critical Features

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yamauchi, Toshinori, Kera, Hiroshi, Kawamoto, Kazuhiko
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918504929689600
author Yamauchi, Toshinori
Kera, Hiroshi
Kawamoto, Kazuhiko
author_facet Yamauchi, Toshinori
Kera, Hiroshi
Kawamoto, Kazuhiko
contents Textual explanations make image classifier decisions transparent by describing the prediction rationale in natural language. Large vision-language models can generate captions but are designed for general visual understanding, not classifier-specific reasoning. Existing zero-shot explanation methods align global image features with language, producing descriptions of what is visible rather than what drives the prediction. We propose TEXTER, which overcomes this limitation by isolating decision-critical features before alignment. TEXTER identifies the neurons contributing to the prediction and emphasizes the features encoded in those neurons -- i.e., the decision-critical features. It then maps these emphasized features into the CLIP feature space to retrieve textual explanations that reflect the model's reasoning. A sparse autoencoder further improves interpretability, particularly for Transformer architectures. Extensive experiments show that TEXTER provides more faithful and interpretable explanations than existing methods. The code is available at \url{https://github.com/tttt-0814/TEXTER}.
format Preprint
id arxiv_https___arxiv_org_abs_2512_07245
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-Shot Textual Explanations via Translating Decision-Critical Features
Yamauchi, Toshinori
Kera, Hiroshi
Kawamoto, Kazuhiko
Computer Vision and Pattern Recognition
Textual explanations make image classifier decisions transparent by describing the prediction rationale in natural language. Large vision-language models can generate captions but are designed for general visual understanding, not classifier-specific reasoning. Existing zero-shot explanation methods align global image features with language, producing descriptions of what is visible rather than what drives the prediction. We propose TEXTER, which overcomes this limitation by isolating decision-critical features before alignment. TEXTER identifies the neurons contributing to the prediction and emphasizes the features encoded in those neurons -- i.e., the decision-critical features. It then maps these emphasized features into the CLIP feature space to retrieve textual explanations that reflect the model's reasoning. A sparse autoencoder further improves interpretability, particularly for Transformer architectures. Extensive experiments show that TEXTER provides more faithful and interpretable explanations than existing methods. The code is available at \url{https://github.com/tttt-0814/TEXTER}.
title Zero-Shot Textual Explanations via Translating Decision-Critical Features
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.07245