Feature compression is the root cause of adversarial fragility in neural network classifiers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Jingchao, Lu, Ziqing, Mudumbai, Raghu, Wu, Xiaodong, Yi, Jirong, Cho, Myung, Xu, Catherine, Xie, Hui, Xu, Weiyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914132994818048
author Gao, Jingchao
Lu, Ziqing
Mudumbai, Raghu
Wu, Xiaodong
Yi, Jirong
Cho, Myung
Xu, Catherine
Xie, Hui
Xu, Weiyu
author_facet Gao, Jingchao
Lu, Ziqing
Mudumbai, Raghu
Wu, Xiaodong
Yi, Jirong
Cho, Myung
Xu, Catherine
Xie, Hui
Xu, Weiyu
contents In this paper, we uniquely study the adversarial robustness of deep neural networks (NN) for classification tasks against that of optimal classifiers. We look at the smallest magnitude of possible additive perturbations that can change a classifier's output. We provide a matrix-theoretic explanation of the adversarial fragility of deep neural networks for classification. In particular, our theoretical results show that a neural network's adversarial robustness can degrade as the input dimension $d$ increases. Analytically, we show that neural networks' adversarial robustness can be only $1/\sqrt{d}$ of the best possible adversarial robustness of optimal classifiers. Our theories match remarkably well with numerical experiments of practically trained NN, including NN for ImageNet images. The matrix-theoretic explanation is consistent with an earlier information-theoretic feature-compression-based explanation for the adversarial fragility of neural networks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_16200
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Feature compression is the root cause of adversarial fragility in neural network classifiers
Gao, Jingchao
Lu, Ziqing
Mudumbai, Raghu
Wu, Xiaodong
Yi, Jirong
Cho, Myung
Xu, Catherine
Xie, Hui
Xu, Weiyu
Machine Learning
Cryptography and Security
Information Theory
Signal Processing
In this paper, we uniquely study the adversarial robustness of deep neural networks (NN) for classification tasks against that of optimal classifiers. We look at the smallest magnitude of possible additive perturbations that can change a classifier's output. We provide a matrix-theoretic explanation of the adversarial fragility of deep neural networks for classification. In particular, our theoretical results show that a neural network's adversarial robustness can degrade as the input dimension $d$ increases. Analytically, we show that neural networks' adversarial robustness can be only $1/\sqrt{d}$ of the best possible adversarial robustness of optimal classifiers. Our theories match remarkably well with numerical experiments of practically trained NN, including NN for ImageNet images. The matrix-theoretic explanation is consistent with an earlier information-theoretic feature-compression-based explanation for the adversarial fragility of neural networks.
title Feature compression is the root cause of adversarial fragility in neural network classifiers
topic Machine Learning
Cryptography and Security
Information Theory
Signal Processing
url https://arxiv.org/abs/2406.16200