Rethinking Robustness: A New Approach to Evaluating Feature Attribution Methods

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kiourti, Panagiota, Singh, Anu, Duraipandian, Preeti, Zhou, Weichao, Li, Wenchao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911306261463040
author Kiourti, Panagiota
Singh, Anu
Duraipandian, Preeti
Zhou, Weichao
Li, Wenchao
author_facet Kiourti, Panagiota
Singh, Anu
Duraipandian, Preeti
Zhou, Weichao
Li, Wenchao
contents This paper studies the robustness of feature attribution methods for deep neural networks. It challenges the current notion of attributional robustness that largely ignores the difference in the model's outputs and introduces a new way of evaluating the robustness of attribution methods. Specifically, we propose a new definition of similar inputs, a new robustness metric, and a novel method based on generative adversarial networks to generate these inputs. In addition, we present a comprehensive evaluation with existing metrics and state-of-the-art attribution methods. Our findings highlight the need for a more objective metric that reveals the weaknesses of an attribution method rather than that of the neural network, thus providing a more accurate evaluation of the robustness of attribution methods.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06665
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Robustness: A New Approach to Evaluating Feature Attribution Methods
Kiourti, Panagiota
Singh, Anu
Duraipandian, Preeti
Zhou, Weichao
Li, Wenchao
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
This paper studies the robustness of feature attribution methods for deep neural networks. It challenges the current notion of attributional robustness that largely ignores the difference in the model's outputs and introduces a new way of evaluating the robustness of attribution methods. Specifically, we propose a new definition of similar inputs, a new robustness metric, and a novel method based on generative adversarial networks to generate these inputs. In addition, we present a comprehensive evaluation with existing metrics and state-of-the-art attribution methods. Our findings highlight the need for a more objective metric that reveals the weaknesses of an attribution method rather than that of the neural network, thus providing a more accurate evaluation of the robustness of attribution methods.
title Rethinking Robustness: A New Approach to Evaluating Feature Attribution Methods
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.06665