A Vulnerability of Attribution Methods Using Pre-Softmax Scores

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lerma, Miguel, Lucas, Mirtha
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909163674664960
author Lerma, Miguel
Lucas, Mirtha
author_facet Lerma, Miguel
Lucas, Mirtha
contents We discuss a vulnerability involving a category of attribution methods used to provide explanations for the outputs of convolutional neural networks working as classifiers. It is known that this type of networks are vulnerable to adversarial attacks, in which imperceptible perturbations of the input may alter the outputs of the model. In contrast, here we focus on effects that small modifications in the model may cause on the attribution method without altering the model outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2307_03305
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Vulnerability of Attribution Methods Using Pre-Softmax Scores
Lerma, Miguel
Lucas, Mirtha
Machine Learning
Artificial Intelligence
68T07
I.2.m
We discuss a vulnerability involving a category of attribution methods used to provide explanations for the outputs of convolutional neural networks working as classifiers. It is known that this type of networks are vulnerable to adversarial attacks, in which imperceptible perturbations of the input may alter the outputs of the model. In contrast, here we focus on effects that small modifications in the model may cause on the attribution method without altering the model outputs.
title A Vulnerability of Attribution Methods Using Pre-Softmax Scores
topic Machine Learning
Artificial Intelligence
68T07
I.2.m
url https://arxiv.org/abs/2307.03305