Classifier Guidance Enhances Diffusion-based Adversarial Purification by Preserving Predictive Information

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Mingkun, Li, Jianing, Chen, Wei, Guo, Jiafeng, Cheng, Xueqi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913465610797056
author Zhang, Mingkun
Li, Jianing
Chen, Wei
Guo, Jiafeng
Cheng, Xueqi
author_facet Zhang, Mingkun
Li, Jianing
Chen, Wei
Guo, Jiafeng
Cheng, Xueqi
contents Adversarial purification is one of the promising approaches to defend neural networks against adversarial attacks. Recently, methods utilizing diffusion probabilistic models have achieved great success for adversarial purification in image classification tasks. However, such methods fall into the dilemma of balancing the needs for noise removal and information preservation. This paper points out that existing adversarial purification methods based on diffusion models gradually lose sample information during the core denoising process, causing occasional label shift in subsequent classification tasks. As a remedy, we suggest to suppress such information loss by introducing guidance from the classifier confidence. Specifically, we propose Classifier-cOnfidence gUided Purification (COUP) algorithm, which purifies adversarial examples while keeping away from the classifier decision boundary. Experimental results show that COUP can achieve better adversarial robustness under strong attack methods.
format Preprint
id arxiv_https___arxiv_org_abs_2408_05900
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Classifier Guidance Enhances Diffusion-based Adversarial Purification by Preserving Predictive Information
Zhang, Mingkun
Li, Jianing
Chen, Wei
Guo, Jiafeng
Cheng, Xueqi
Computer Vision and Pattern Recognition
Adversarial purification is one of the promising approaches to defend neural networks against adversarial attacks. Recently, methods utilizing diffusion probabilistic models have achieved great success for adversarial purification in image classification tasks. However, such methods fall into the dilemma of balancing the needs for noise removal and information preservation. This paper points out that existing adversarial purification methods based on diffusion models gradually lose sample information during the core denoising process, causing occasional label shift in subsequent classification tasks. As a remedy, we suggest to suppress such information loss by introducing guidance from the classifier confidence. Specifically, we propose Classifier-cOnfidence gUided Purification (COUP) algorithm, which purifies adversarial examples while keeping away from the classifier decision boundary. Experimental results show that COUP can achieve better adversarial robustness under strong attack methods.
title Classifier Guidance Enhances Diffusion-based Adversarial Purification by Preserving Predictive Information
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.05900