Attribution for Enhanced Explanation with Transferable Adversarial eXploration

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhu, Zhiyu, Zhang, Jiayu, Jin, Zhibo, Chen, Huaming, Zhou, Jianlong, Chen, Fang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929648872456192
author Zhu, Zhiyu
Zhang, Jiayu
Jin, Zhibo
Chen, Huaming
Zhou, Jianlong
Chen, Fang
author_facet Zhu, Zhiyu
Zhang, Jiayu
Jin, Zhibo
Chen, Huaming
Zhou, Jianlong
Chen, Fang
contents The interpretability of deep neural networks is crucial for understanding model decisions in various applications, including computer vision. AttEXplore++, an advanced framework built upon AttEXplore, enhances attribution by incorporating transferable adversarial attack methods such as MIG and GRA, significantly improving the accuracy and robustness of model explanations. We conduct extensive experiments on five models, including CNNs (Inception-v3, ResNet-50, VGG16) and vision transformers (MaxViT-T, ViT-B/16), using the ImageNet dataset. Our method achieves an average performance improvement of 7.57\% over AttEXplore and 32.62\% compared to other state-of-the-art interpretability algorithms. Using insertion and deletion scores as evaluation metrics, we show that adversarial transferability plays a vital role in enhancing attribution results. Furthermore, we explore the impact of randomness, perturbation rate, noise amplitude, and diversity probability on attribution performance, demonstrating that AttEXplore++ provides more stable and reliable explanations across various models. We release our code at: https://anonymous.4open.science/r/ATTEXPLOREP-8435/
format Preprint
id arxiv_https___arxiv_org_abs_2412_19523
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Attribution for Enhanced Explanation with Transferable Adversarial eXploration
Zhu, Zhiyu
Zhang, Jiayu
Jin, Zhibo
Chen, Huaming
Zhou, Jianlong
Chen, Fang
Artificial Intelligence
Computer Vision and Pattern Recognition
The interpretability of deep neural networks is crucial for understanding model decisions in various applications, including computer vision. AttEXplore++, an advanced framework built upon AttEXplore, enhances attribution by incorporating transferable adversarial attack methods such as MIG and GRA, significantly improving the accuracy and robustness of model explanations. We conduct extensive experiments on five models, including CNNs (Inception-v3, ResNet-50, VGG16) and vision transformers (MaxViT-T, ViT-B/16), using the ImageNet dataset. Our method achieves an average performance improvement of 7.57\% over AttEXplore and 32.62\% compared to other state-of-the-art interpretability algorithms. Using insertion and deletion scores as evaluation metrics, we show that adversarial transferability plays a vital role in enhancing attribution results. Furthermore, we explore the impact of randomness, perturbation rate, noise amplitude, and diversity probability on attribution performance, demonstrating that AttEXplore++ provides more stable and reliable explanations across various models. We release our code at: https://anonymous.4open.science/r/ATTEXPLOREP-8435/
title Attribution for Enhanced Explanation with Transferable Adversarial eXploration
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.19523