Large-Scale Dataset Pruning in Adversarial Training through Data Importance Extrapolation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nieth, Björn, Altstidl, Thomas, Schwinn, Leo, Eskofier, Björn
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929417389867008
author Nieth, Björn
Altstidl, Thomas
Schwinn, Leo
Eskofier, Björn
author_facet Nieth, Björn
Altstidl, Thomas
Schwinn, Leo
Eskofier, Björn
contents Their vulnerability to small, imperceptible attacks limits the adoption of deep learning models to real-world systems. Adversarial training has proven to be one of the most promising strategies against these attacks, at the expense of a substantial increase in training time. With the ongoing trend of integrating large-scale synthetic data this is only expected to increase even further. Thus, the need for data-centric approaches that reduce the number of training samples while maintaining accuracy and robustness arises. While data pruning and active learning are prominent research topics in deep learning, they are as of now largely unexplored in the adversarial training literature. We address this gap and propose a new data pruning strategy based on extrapolating data importance scores from a small set of data to a larger set. In an empirical evaluation, we demonstrate that extrapolation-based pruning can efficiently reduce dataset size while maintaining robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13283
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large-Scale Dataset Pruning in Adversarial Training through Data Importance Extrapolation
Nieth, Björn
Altstidl, Thomas
Schwinn, Leo
Eskofier, Björn
Machine Learning
Their vulnerability to small, imperceptible attacks limits the adoption of deep learning models to real-world systems. Adversarial training has proven to be one of the most promising strategies against these attacks, at the expense of a substantial increase in training time. With the ongoing trend of integrating large-scale synthetic data this is only expected to increase even further. Thus, the need for data-centric approaches that reduce the number of training samples while maintaining accuracy and robustness arises. While data pruning and active learning are prominent research topics in deep learning, they are as of now largely unexplored in the adversarial training literature. We address this gap and propose a new data pruning strategy based on extrapolating data importance scores from a small set of data to a larger set. In an empirical evaluation, we demonstrate that extrapolation-based pruning can efficiently reduce dataset size while maintaining robustness.
title Large-Scale Dataset Pruning in Adversarial Training through Data Importance Extrapolation
topic Machine Learning
url https://arxiv.org/abs/2406.13283