Exploring the Interplay of Interpretability and Robustness in Deep Neural Networks: A Saliency-guided Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guesmi, Amira, Aswani, Nishant Suresh, Shafique, Muhammad
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914790441484288
author Guesmi, Amira
Aswani, Nishant Suresh
Shafique, Muhammad
author_facet Guesmi, Amira
Aswani, Nishant Suresh
Shafique, Muhammad
contents Adversarial attacks pose a significant challenge to deploying deep learning models in safety-critical applications. Maintaining model robustness while ensuring interpretability is vital for fostering trust and comprehension in these models. This study investigates the impact of Saliency-guided Training (SGT) on model robustness, a technique aimed at improving the clarity of saliency maps to deepen understanding of the model's decision-making process. Experiments were conducted on standard benchmark datasets using various deep learning architectures trained with and without SGT. Findings demonstrate that SGT enhances both model robustness and interpretability. Additionally, we propose a novel approach combining SGT with standard adversarial training to achieve even greater robustness while preserving saliency map quality. Our strategy is grounded in the assumption that preserving salient features crucial for correctly classifying adversarial examples enhances model robustness, while masking non-relevant features improves interpretability. Our technique yields significant gains, achieving a 35\% and 20\% improvement in robustness against PGD attack with noise magnitudes of $0.2$ and $0.02$ for the MNIST and CIFAR-10 datasets, respectively, while producing high-quality saliency maps.
format Preprint
id arxiv_https___arxiv_org_abs_2405_06278
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring the Interplay of Interpretability and Robustness in Deep Neural Networks: A Saliency-guided Approach
Guesmi, Amira
Aswani, Nishant Suresh
Shafique, Muhammad
Computer Vision and Pattern Recognition
Cryptography and Security
Adversarial attacks pose a significant challenge to deploying deep learning models in safety-critical applications. Maintaining model robustness while ensuring interpretability is vital for fostering trust and comprehension in these models. This study investigates the impact of Saliency-guided Training (SGT) on model robustness, a technique aimed at improving the clarity of saliency maps to deepen understanding of the model's decision-making process. Experiments were conducted on standard benchmark datasets using various deep learning architectures trained with and without SGT. Findings demonstrate that SGT enhances both model robustness and interpretability. Additionally, we propose a novel approach combining SGT with standard adversarial training to achieve even greater robustness while preserving saliency map quality. Our strategy is grounded in the assumption that preserving salient features crucial for correctly classifying adversarial examples enhances model robustness, while masking non-relevant features improves interpretability. Our technique yields significant gains, achieving a 35\% and 20\% improvement in robustness against PGD attack with noise magnitudes of $0.2$ and $0.02$ for the MNIST and CIFAR-10 datasets, respectively, while producing high-quality saliency maps.
title Exploring the Interplay of Interpretability and Robustness in Deep Neural Networks: A Saliency-guided Approach
topic Computer Vision and Pattern Recognition
Cryptography and Security
url https://arxiv.org/abs/2405.06278