Annealing Self-Distillation Rectification Improves Adversarial Training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Yu-Yu, Wang, Hung-Jui, Chen, Shang-Tse
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909168352362496
author Wu, Yu-Yu
Wang, Hung-Jui
Chen, Shang-Tse
author_facet Wu, Yu-Yu
Wang, Hung-Jui
Chen, Shang-Tse
contents In standard adversarial training, models are optimized to fit one-hot labels within allowable adversarial perturbation budgets. However, the ignorance of underlying distribution shifts brought by perturbations causes the problem of robust overfitting. To address this issue and enhance adversarial robustness, we analyze the characteristics of robust models and identify that robust models tend to produce smoother and well-calibrated outputs. Based on the observation, we propose a simple yet effective method, Annealing Self-Distillation Rectification (ADR), which generates soft labels as a better guidance mechanism that accurately reflects the distribution shift under attack during adversarial training. By utilizing ADR, we can obtain rectified distributions that significantly improve model robustness without the need for pre-trained models or extensive extra computation. Moreover, our method facilitates seamless plug-and-play integration with other adversarial training techniques by replacing the hard labels in their objectives. We demonstrate the efficacy of ADR through extensive experiments and strong performances across datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2305_12118
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Annealing Self-Distillation Rectification Improves Adversarial Training
Wu, Yu-Yu
Wang, Hung-Jui
Chen, Shang-Tse
Machine Learning
Artificial Intelligence
In standard adversarial training, models are optimized to fit one-hot labels within allowable adversarial perturbation budgets. However, the ignorance of underlying distribution shifts brought by perturbations causes the problem of robust overfitting. To address this issue and enhance adversarial robustness, we analyze the characteristics of robust models and identify that robust models tend to produce smoother and well-calibrated outputs. Based on the observation, we propose a simple yet effective method, Annealing Self-Distillation Rectification (ADR), which generates soft labels as a better guidance mechanism that accurately reflects the distribution shift under attack during adversarial training. By utilizing ADR, we can obtain rectified distributions that significantly improve model robustness without the need for pre-trained models or extensive extra computation. Moreover, our method facilitates seamless plug-and-play integration with other adversarial training techniques by replacing the hard labels in their objectives. We demonstrate the efficacy of ADR through extensive experiments and strong performances across datasets.
title Annealing Self-Distillation Rectification Improves Adversarial Training
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2305.12118