Wide Two-Layer Networks can Learn from Adversarial Perturbations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumano, Soichiro, Kera, Hiroshi, Yamasaki, Toshihiko
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929681423400960
author Kumano, Soichiro
Kera, Hiroshi
Yamasaki, Toshihiko
author_facet Kumano, Soichiro
Kera, Hiroshi
Yamasaki, Toshihiko
contents Adversarial examples have raised several open questions, such as why they can deceive classifiers and transfer between different models. A prevailing hypothesis to explain these phenomena suggests that adversarial perturbations appear as random noise but contain class-specific features. This hypothesis is supported by the success of perturbation learning, where classifiers trained solely on adversarial examples and the corresponding incorrect labels generalize well to correctly labeled test data. Although this hypothesis and perturbation learning are effective in explaining intriguing properties of adversarial examples, their solid theoretical foundation is limited. In this study, we theoretically explain the counterintuitive success of perturbation learning. We assume wide two-layer networks and the results hold for any data distribution. We prove that adversarial perturbations contain sufficient class-specific features for networks to generalize from them. Moreover, the predictions of classifiers trained on mislabeled adversarial examples coincide with those of classifiers trained on correctly labeled clean samples. The code is available at https://github.com/s-kumano/perturbation-learning.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23677
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Wide Two-Layer Networks can Learn from Adversarial Perturbations
Kumano, Soichiro
Kera, Hiroshi
Yamasaki, Toshihiko
Machine Learning
Computer Vision and Pattern Recognition
Adversarial examples have raised several open questions, such as why they can deceive classifiers and transfer between different models. A prevailing hypothesis to explain these phenomena suggests that adversarial perturbations appear as random noise but contain class-specific features. This hypothesis is supported by the success of perturbation learning, where classifiers trained solely on adversarial examples and the corresponding incorrect labels generalize well to correctly labeled test data. Although this hypothesis and perturbation learning are effective in explaining intriguing properties of adversarial examples, their solid theoretical foundation is limited. In this study, we theoretically explain the counterintuitive success of perturbation learning. We assume wide two-layer networks and the results hold for any data distribution. We prove that adversarial perturbations contain sufficient class-specific features for networks to generalize from them. Moreover, the predictions of classifiers trained on mislabeled adversarial examples coincide with those of classifiers trained on correctly labeled clean samples. The code is available at https://github.com/s-kumano/perturbation-learning.
title Wide Two-Layer Networks can Learn from Adversarial Perturbations
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.23677