On the Impact of Hard Adversarial Instances on Overfitting in Adversarial Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Chen, Huang, Zhichao, Salzmann, Mathieu, Zhang, Tong, Süsstrunk, Sabine
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929634197635072
author Liu, Chen
Huang, Zhichao
Salzmann, Mathieu
Zhang, Tong
Süsstrunk, Sabine
author_facet Liu, Chen
Huang, Zhichao
Salzmann, Mathieu
Zhang, Tong
Süsstrunk, Sabine
contents Adversarial training is a popular method to robustify models against adversarial attacks. However, it exhibits much more severe overfitting than training on clean inputs. In this work, we investigate this phenomenon from the perspective of training instances, i.e., training input-target pairs. Based on a quantitative metric measuring the relative difficulty of an instance in the training set, we analyze the model's behavior on training instances of different difficulty levels. This lets us demonstrate that the decay in generalization performance of adversarial training is a result of fitting hard adversarial instances. We theoretically verify our observations for both linear and general nonlinear models, proving that models trained on hard instances have worse generalization performance than ones trained on easy instances, and that this generalization gap increases with the size of the adversarial budget. Finally, we investigate solutions to mitigate adversarial overfitting in several scenarios, including fast adversarial training and fine-tuning a pretrained model with additional data. Our results demonstrate that using training data adaptively improves the model's robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2112_07324
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle On the Impact of Hard Adversarial Instances on Overfitting in Adversarial Training
Liu, Chen
Huang, Zhichao
Salzmann, Mathieu
Zhang, Tong
Süsstrunk, Sabine
Machine Learning
Adversarial training is a popular method to robustify models against adversarial attacks. However, it exhibits much more severe overfitting than training on clean inputs. In this work, we investigate this phenomenon from the perspective of training instances, i.e., training input-target pairs. Based on a quantitative metric measuring the relative difficulty of an instance in the training set, we analyze the model's behavior on training instances of different difficulty levels. This lets us demonstrate that the decay in generalization performance of adversarial training is a result of fitting hard adversarial instances. We theoretically verify our observations for both linear and general nonlinear models, proving that models trained on hard instances have worse generalization performance than ones trained on easy instances, and that this generalization gap increases with the size of the adversarial budget. Finally, we investigate solutions to mitigate adversarial overfitting in several scenarios, including fast adversarial training and fine-tuning a pretrained model with additional data. Our results demonstrate that using training data adaptively improves the model's robustness.
title On the Impact of Hard Adversarial Instances on Overfitting in Adversarial Training
topic Machine Learning
url https://arxiv.org/abs/2112.07324