Adversarial Samples Are Not Created Equal

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Crawford, Jennifer, Khanna, Amol, Lu, Fred, Wagoner, Amy R., Biderman, Stella, Nguyen, Andre T., Raff, Edward
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909979936555008
author Crawford, Jennifer
Khanna, Amol
Lu, Fred
Wagoner, Amy R.
Biderman, Stella
Nguyen, Andre T.
Raff, Edward
author_facet Crawford, Jennifer
Khanna, Amol
Lu, Fred
Wagoner, Amy R.
Biderman, Stella
Nguyen, Andre T.
Raff, Edward
contents Over the past decade, numerous theories have been proposed to explain the widespread vulnerability of deep neural networks to adversarial evasion attacks. Among these, the theory of non-robust features proposed by Ilyas et al. has been widely accepted, showing that brittle but predictive features of the data distribution can be directly exploited by attackers. However, this theory overlooks adversarial samples that do not directly utilize these features. In this work, we advocate that these two kinds of samples - those which use use brittle but predictive features and those that do not - comprise two types of adversarial weaknesses and should be differentiated when evaluating adversarial robustness. For this purpose, we propose an ensemble-based metric to measure the manipulation of non-robust features by adversarial perturbations and use this metric to analyze the makeup of adversarial samples generated by attackers. This new perspective also allows us to re-examine multiple phenomena, including the impact of sharpness-aware minimization on adversarial robustness and the robustness gap observed between adversarially training and standard training on robust datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2601_00577
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Adversarial Samples Are Not Created Equal
Crawford, Jennifer
Khanna, Amol
Lu, Fred
Wagoner, Amy R.
Biderman, Stella
Nguyen, Andre T.
Raff, Edward
Machine Learning
Over the past decade, numerous theories have been proposed to explain the widespread vulnerability of deep neural networks to adversarial evasion attacks. Among these, the theory of non-robust features proposed by Ilyas et al. has been widely accepted, showing that brittle but predictive features of the data distribution can be directly exploited by attackers. However, this theory overlooks adversarial samples that do not directly utilize these features. In this work, we advocate that these two kinds of samples - those which use use brittle but predictive features and those that do not - comprise two types of adversarial weaknesses and should be differentiated when evaluating adversarial robustness. For this purpose, we propose an ensemble-based metric to measure the manipulation of non-robust features by adversarial perturbations and use this metric to analyze the makeup of adversarial samples generated by attackers. This new perspective also allows us to re-examine multiple phenomena, including the impact of sharpness-aware minimization on adversarial robustness and the robustness gap observed between adversarially training and standard training on robust datasets.
title Adversarial Samples Are Not Created Equal
topic Machine Learning
url https://arxiv.org/abs/2601.00577