Enhancing Interpretability for Vision Models via Shapley Value Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fan, Kanglong, Yang, Yunqiao, Ma, Chen
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914204019064832
author Fan, Kanglong
Yang, Yunqiao
Ma, Chen
author_facet Fan, Kanglong
Yang, Yunqiao
Ma, Chen
contents Deep neural networks have demonstrated remarkable performance across various domains, yet their decision-making processes remain opaque. Although many explanation methods are dedicated to bringing the obscurity of DNNs to light, they exhibit significant limitations: post-hoc explanation methods often struggle to faithfully reflect model behaviors, while self-explaining neural networks sacrifice performance and compatibility due to their specialized architectural designs. To address these challenges, we propose a novel self-explaining framework that integrates Shapley value estimation as an auxiliary task during training, which achieves two key advancements: 1) a fair allocation of the model prediction scores to image patches, ensuring explanations inherently align with the model's decision logic, and 2) enhanced interpretability with minor structural modifications, preserving model performance and compatibility. Extensive experiments on multiple benchmarks demonstrate that our method achieves state-of-the-art interpretability.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14354
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Interpretability for Vision Models via Shapley Value Optimization
Fan, Kanglong
Yang, Yunqiao
Ma, Chen
Computer Vision and Pattern Recognition
Artificial Intelligence
Deep neural networks have demonstrated remarkable performance across various domains, yet their decision-making processes remain opaque. Although many explanation methods are dedicated to bringing the obscurity of DNNs to light, they exhibit significant limitations: post-hoc explanation methods often struggle to faithfully reflect model behaviors, while self-explaining neural networks sacrifice performance and compatibility due to their specialized architectural designs. To address these challenges, we propose a novel self-explaining framework that integrates Shapley value estimation as an auxiliary task during training, which achieves two key advancements: 1) a fair allocation of the model prediction scores to image patches, ensuring explanations inherently align with the model's decision logic, and 2) enhanced interpretability with minor structural modifications, preserving model performance and compatibility. Extensive experiments on multiple benchmarks demonstrate that our method achieves state-of-the-art interpretability.
title Enhancing Interpretability for Vision Models via Shapley Value Optimization
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.14354