CausalDiff: Causality-Inspired Disentanglement via Diffusion Model for Adversarial Defense

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Mingkun, Bi, Keping, Chen, Wei, Chen, Quanrun, Guo, Jiafeng, Cheng, Xueqi
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913706463461376
author Zhang, Mingkun
Bi, Keping
Chen, Wei
Chen, Quanrun
Guo, Jiafeng
Cheng, Xueqi
author_facet Zhang, Mingkun
Bi, Keping
Chen, Wei
Chen, Quanrun
Guo, Jiafeng
Cheng, Xueqi
contents Despite ongoing efforts to defend neural classifiers from adversarial attacks, they remain vulnerable, especially to unseen attacks. In contrast, humans are difficult to be cheated by subtle manipulations, since we make judgments only based on essential factors. Inspired by this observation, we attempt to model label generation with essential label-causative factors and incorporate label-non-causative factors to assist data generation. For an adversarial example, we aim to discriminate the perturbations as non-causative factors and make predictions only based on the label-causative factors. Concretely, we propose a casual diffusion model (CausalDiff) that adapts diffusion models for conditional data generation and disentangles the two types of casual factors by learning towards a novel casual information bottleneck objective. Empirically, CausalDiff has significantly outperformed state-of-the-art defense methods on various unseen attacks, achieving an average robustness of 86.39% (+4.01%) on CIFAR-10, 56.25% (+3.13%) on CIFAR-100, and 82.62% (+4.93%) on GTSRB (German Traffic Sign Recognition Benchmark). The code is available at https://github.com/CAS-AISafetyBasicResearchGroup/CausalDiff.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23091
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CausalDiff: Causality-Inspired Disentanglement via Diffusion Model for Adversarial Defense
Zhang, Mingkun
Bi, Keping
Chen, Wei
Chen, Quanrun
Guo, Jiafeng
Cheng, Xueqi
Computer Vision and Pattern Recognition
Despite ongoing efforts to defend neural classifiers from adversarial attacks, they remain vulnerable, especially to unseen attacks. In contrast, humans are difficult to be cheated by subtle manipulations, since we make judgments only based on essential factors. Inspired by this observation, we attempt to model label generation with essential label-causative factors and incorporate label-non-causative factors to assist data generation. For an adversarial example, we aim to discriminate the perturbations as non-causative factors and make predictions only based on the label-causative factors. Concretely, we propose a casual diffusion model (CausalDiff) that adapts diffusion models for conditional data generation and disentangles the two types of casual factors by learning towards a novel casual information bottleneck objective. Empirically, CausalDiff has significantly outperformed state-of-the-art defense methods on various unseen attacks, achieving an average robustness of 86.39% (+4.01%) on CIFAR-10, 56.25% (+3.13%) on CIFAR-100, and 82.62% (+4.93%) on GTSRB (German Traffic Sign Recognition Benchmark). The code is available at https://github.com/CAS-AISafetyBasicResearchGroup/CausalDiff.
title CausalDiff: Causality-Inspired Disentanglement via Diffusion Model for Adversarial Defense
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.23091