A Backdoor-based Explainable AI Benchmark for High Fidelity Evaluation of Attributions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Peiyu, Akhtar, Naveed, Jiang, Jiantong, Mian, Ajmal
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918151818575872
author Yang, Peiyu
Akhtar, Naveed
Jiang, Jiantong
Mian, Ajmal
author_facet Yang, Peiyu
Akhtar, Naveed
Jiang, Jiantong
Mian, Ajmal
contents Attribution methods compute importance scores for input features to explain model predictions. However, assessing the faithfulness of these methods remains challenging due to the absence of attribution ground truth to model predictions. In this work, we first identify a set of fidelity criteria that reliable benchmarks for attribution methods are expected to fulfill, thereby facilitating a systematic assessment of attribution benchmarks. Next, we introduce a Backdoor-based eXplainable AI benchmark (BackX) that adheres to the desired fidelity criteria. We theoretically establish the superiority of our approach over the existing benchmarks for well-founded attribution evaluation. With extensive analysis, we further establish a standardized evaluation setup that mitigates confounding factors such as post-processing techniques and explained predictions, thereby ensuring a fair and consistent benchmarking. This setup is ultimately employed for a comprehensive comparison of existing methods using BackX. Finally, our analysis also offers insights into defending against neural Trojans by utilizing the attributions.
format Preprint
id arxiv_https___arxiv_org_abs_2405_02344
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Backdoor-based Explainable AI Benchmark for High Fidelity Evaluation of Attributions
Yang, Peiyu
Akhtar, Naveed
Jiang, Jiantong
Mian, Ajmal
Cryptography and Security
Artificial Intelligence
Machine Learning
Attribution methods compute importance scores for input features to explain model predictions. However, assessing the faithfulness of these methods remains challenging due to the absence of attribution ground truth to model predictions. In this work, we first identify a set of fidelity criteria that reliable benchmarks for attribution methods are expected to fulfill, thereby facilitating a systematic assessment of attribution benchmarks. Next, we introduce a Backdoor-based eXplainable AI benchmark (BackX) that adheres to the desired fidelity criteria. We theoretically establish the superiority of our approach over the existing benchmarks for well-founded attribution evaluation. With extensive analysis, we further establish a standardized evaluation setup that mitigates confounding factors such as post-processing techniques and explained predictions, thereby ensuring a fair and consistent benchmarking. This setup is ultimately employed for a comprehensive comparison of existing methods using BackX. Finally, our analysis also offers insights into defending against neural Trojans by utilizing the attributions.
title A Backdoor-based Explainable AI Benchmark for High Fidelity Evaluation of Attributions
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2405.02344