Robustness Feature Adapter for Efficient Adversarial Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Quanwei, Guo, Jun, Wang, Wei, Wang, Yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908500848803840
author Wu, Quanwei
Guo, Jun
Wang, Wei
Wang, Yi
author_facet Wu, Quanwei
Guo, Jun
Wang, Wei
Wang, Yi
contents Adversarial training (AT) with projected gradient descent is the most popular method to improve model robustness under adversarial attacks. However, computational overheads become prohibitively large when AT is applied to large backbone models. AT is also known to have the issue of robust overfitting. This paper contributes to solving both problems simultaneously towards building more trustworthy foundation models. In particular, we propose a new adapter-based approach for efficient AT directly in the feature space. We show that the proposed adapter-based approach can improve the inner-loop convergence quality by eliminating robust overfitting. As a result, it significantly increases computational efficiency and improves model accuracy by generalizing adversarial robustness to unseen attacks. We demonstrate the effectiveness of the new adapter-based approach in different backbone architectures and in AT at scale.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17680
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robustness Feature Adapter for Efficient Adversarial Training
Wu, Quanwei
Guo, Jun
Wang, Wei
Wang, Yi
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.6
Adversarial training (AT) with projected gradient descent is the most popular method to improve model robustness under adversarial attacks. However, computational overheads become prohibitively large when AT is applied to large backbone models. AT is also known to have the issue of robust overfitting. This paper contributes to solving both problems simultaneously towards building more trustworthy foundation models. In particular, we propose a new adapter-based approach for efficient AT directly in the feature space. We show that the proposed adapter-based approach can improve the inner-loop convergence quality by eliminating robust overfitting. As a result, it significantly increases computational efficiency and improves model accuracy by generalizing adversarial robustness to unseen attacks. We demonstrate the effectiveness of the new adapter-based approach in different backbone architectures and in AT at scale.
title Robustness Feature Adapter for Efficient Adversarial Training
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.6
url https://arxiv.org/abs/2508.17680