When and How to Fool Explainable Models (and Humans) with Adversarial Examples

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vadillo, Jon, Santana, Roberto, Lozano, Jose A.
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910828750438400
author Vadillo, Jon
Santana, Roberto
Lozano, Jose A.
author_facet Vadillo, Jon
Santana, Roberto
Lozano, Jose A.
contents Reliable deployment of machine learning models such as neural networks continues to be challenging due to several limitations. Some of the main shortcomings are the lack of interpretability and the lack of robustness against adversarial examples or out-of-distribution inputs. In this exploratory review, we explore the possibilities and limits of adversarial attacks for explainable machine learning models. First, we extend the notion of adversarial examples to fit in explainable machine learning scenarios, in which the inputs, the output classifications and the explanations of the model's decisions are assessed by humans. Next, we propose a comprehensive framework to study whether (and how) adversarial examples can be generated for explainable models under human assessment, introducing and illustrating novel attack paradigms. In particular, our framework considers a wide range of relevant yet often ignored factors such as the type of problem, the user expertise or the objective of the explanations, in order to identify the attack strategies that should be adopted in each scenario to successfully deceive the model (and the human). The intention of these contributions is to serve as a basis for a more rigorous and realistic study of adversarial examples in the field of explainable machine learning.
format Preprint
id arxiv_https___arxiv_org_abs_2107_01943
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle When and How to Fool Explainable Models (and Humans) with Adversarial Examples
Vadillo, Jon
Santana, Roberto
Lozano, Jose A.
Machine Learning
Cryptography and Security
Reliable deployment of machine learning models such as neural networks continues to be challenging due to several limitations. Some of the main shortcomings are the lack of interpretability and the lack of robustness against adversarial examples or out-of-distribution inputs. In this exploratory review, we explore the possibilities and limits of adversarial attacks for explainable machine learning models. First, we extend the notion of adversarial examples to fit in explainable machine learning scenarios, in which the inputs, the output classifications and the explanations of the model's decisions are assessed by humans. Next, we propose a comprehensive framework to study whether (and how) adversarial examples can be generated for explainable models under human assessment, introducing and illustrating novel attack paradigms. In particular, our framework considers a wide range of relevant yet often ignored factors such as the type of problem, the user expertise or the objective of the explanations, in order to identify the attack strategies that should be adopted in each scenario to successfully deceive the model (and the human). The intention of these contributions is to serve as a basis for a more rigorous and realistic study of adversarial examples in the field of explainable machine learning.
title When and How to Fool Explainable Models (and Humans) with Adversarial Examples
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2107.01943