Prediction Inconsistency Helps Achieve Generalizable Detection of Adversarial Examples

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Han, Sicong, Lin, Chenhao, Zhao, Zhengyu, Wang, Xiyuan, He, Xinlei, Li, Qian, Wang, Cong, Wang, Qian, Shen, Chao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915324731850752
author Han, Sicong
Lin, Chenhao
Zhao, Zhengyu
Wang, Xiyuan
He, Xinlei
Li, Qian
Wang, Cong
Wang, Qian
Shen, Chao
author_facet Han, Sicong
Lin, Chenhao
Zhao, Zhengyu
Wang, Xiyuan
He, Xinlei
Li, Qian
Wang, Cong
Wang, Qian
Shen, Chao
contents Adversarial detection protects models from adversarial attacks by refusing suspicious test samples. However, current detection methods often suffer from weak generalization: their effectiveness tends to degrade significantly when applied to adversarially trained models rather than naturally trained ones, and they generally struggle to achieve consistent effectiveness across both white-box and black-box attack settings. In this work, we observe that an auxiliary model, differing from the primary model in training strategy or model architecture, tends to assign low confidence to the primary model's predictions on adversarial examples (AEs), while preserving high confidence on normal examples (NEs). Based on this discovery, we propose Prediction Inconsistency Detector (PID), a lightweight and generalizable detection framework to distinguish AEs from NEs by capturing the prediction inconsistency between the primal and auxiliary models. PID is compatible with both naturally and adversarially trained primal models and outperforms four detection methods across 3 white-box, 3 black-box, and 1 mixed adversarial attacks. Specifically, PID achieves average AUC scores of 99.29\% and 99.30\% on CIFAR-10 when the primal model is naturally and adversarially trained, respectively, and 98.31% and 96.81% on ImageNet under the same conditions, outperforming existing SOTAs by 4.70%$\sim$25.46%.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03765
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prediction Inconsistency Helps Achieve Generalizable Detection of Adversarial Examples
Han, Sicong
Lin, Chenhao
Zhao, Zhengyu
Wang, Xiyuan
He, Xinlei
Li, Qian
Wang, Cong
Wang, Qian
Shen, Chao
Cryptography and Security
Adversarial detection protects models from adversarial attacks by refusing suspicious test samples. However, current detection methods often suffer from weak generalization: their effectiveness tends to degrade significantly when applied to adversarially trained models rather than naturally trained ones, and they generally struggle to achieve consistent effectiveness across both white-box and black-box attack settings. In this work, we observe that an auxiliary model, differing from the primary model in training strategy or model architecture, tends to assign low confidence to the primary model's predictions on adversarial examples (AEs), while preserving high confidence on normal examples (NEs). Based on this discovery, we propose Prediction Inconsistency Detector (PID), a lightweight and generalizable detection framework to distinguish AEs from NEs by capturing the prediction inconsistency between the primal and auxiliary models. PID is compatible with both naturally and adversarially trained primal models and outperforms four detection methods across 3 white-box, 3 black-box, and 1 mixed adversarial attacks. Specifically, PID achieves average AUC scores of 99.29\% and 99.30\% on CIFAR-10 when the primal model is naturally and adversarially trained, respectively, and 98.31% and 96.81% on ImageNet under the same conditions, outperforming existing SOTAs by 4.70%$\sim$25.46%.
title Prediction Inconsistency Helps Achieve Generalizable Detection of Adversarial Examples
topic Cryptography and Security
url https://arxiv.org/abs/2506.03765