Zero-Shot Faithful Textual Explanations via Directional-Derivative Influence on Predictions

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yamauchi, Toshinori, Kera, Hiroshi, Kawamoto, Kazuhiko
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910226059362304
author Yamauchi, Toshinori
Kera, Hiroshi
Kawamoto, Kazuhiko
author_facet Yamauchi, Toshinori
Kera, Hiroshi
Kawamoto, Kazuhiko
contents Zero-shot textual explanations aim to make image classifiers more transparent by probing their internal representations, without relying on task-specific supervision or LVLMs. However, existing methods often miss the features that truly drive the prediction, resulting in limited \textit{faithfulness} to the evidence underlying the model's decision. To address this, we propose FaithTrace. Motivated by the idea that faithful explanations should describe concepts that strongly influence the prediction, FaithTrace directly measures how much the representation induced by the explanation changes the class logit. We introduce an influence score, computed as the directional derivative of the class logit along the text-induced direction in the classifier's feature space, and use it as a proxy for faithfulness. Moreover, we extend this influence score into quantitative evaluation metrics, helping fill the gap in faithfulness evaluation for textual explanations. Experiments show that FaithTrace yields more faithful explanations than baselines, facilitating a more accurate understanding of the model. The code will be publicly released.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16877
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Zero-Shot Faithful Textual Explanations via Directional-Derivative Influence on Predictions
Yamauchi, Toshinori
Kera, Hiroshi
Kawamoto, Kazuhiko
Computer Vision and Pattern Recognition
Zero-shot textual explanations aim to make image classifiers more transparent by probing their internal representations, without relying on task-specific supervision or LVLMs. However, existing methods often miss the features that truly drive the prediction, resulting in limited \textit{faithfulness} to the evidence underlying the model's decision. To address this, we propose FaithTrace. Motivated by the idea that faithful explanations should describe concepts that strongly influence the prediction, FaithTrace directly measures how much the representation induced by the explanation changes the class logit. We introduce an influence score, computed as the directional derivative of the class logit along the text-induced direction in the classifier's feature space, and use it as a proxy for faithfulness. Moreover, we extend this influence score into quantitative evaluation metrics, helping fill the gap in faithfulness evaluation for textual explanations. Experiments show that FaithTrace yields more faithful explanations than baselines, facilitating a more accurate understanding of the model. The code will be publicly released.
title Zero-Shot Faithful Textual Explanations via Directional-Derivative Influence on Predictions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.16877