Adversarial Examples are Misaligned in Diffusion Model Manifolds

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lorenz, Peter, Durall, Ricard, Keuper, Janis
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910370337128448
author Lorenz, Peter
Durall, Ricard
Keuper, Janis
author_facet Lorenz, Peter
Durall, Ricard
Keuper, Janis
contents In recent years, diffusion models (DMs) have drawn significant attention for their success in approximating data distributions, yielding state-of-the-art generative results. Nevertheless, the versatility of these models extends beyond their generative capabilities to encompass various vision applications, such as image inpainting, segmentation, adversarial robustness, among others. This study is dedicated to the investigation of adversarial attacks through the lens of diffusion models. However, our objective does not involve enhancing the adversarial robustness of image classifiers. Instead, our focus lies in utilizing the diffusion model to detect and analyze the anomalies introduced by these attacks on images. To that end, we systematically examine the alignment of the distributions of adversarial examples when subjected to the process of transformation using diffusion models. The efficacy of this approach is assessed across CIFAR-10 and ImageNet datasets, including varying image sizes in the latter. The results demonstrate a notable capacity to discriminate effectively between benign and attacked images, providing compelling evidence that adversarial instances do not align with the learned manifold of the DMs.
format Preprint
id arxiv_https___arxiv_org_abs_2401_06637
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adversarial Examples are Misaligned in Diffusion Model Manifolds
Lorenz, Peter
Durall, Ricard
Keuper, Janis
Computer Vision and Pattern Recognition
Cryptography and Security
In recent years, diffusion models (DMs) have drawn significant attention for their success in approximating data distributions, yielding state-of-the-art generative results. Nevertheless, the versatility of these models extends beyond their generative capabilities to encompass various vision applications, such as image inpainting, segmentation, adversarial robustness, among others. This study is dedicated to the investigation of adversarial attacks through the lens of diffusion models. However, our objective does not involve enhancing the adversarial robustness of image classifiers. Instead, our focus lies in utilizing the diffusion model to detect and analyze the anomalies introduced by these attacks on images. To that end, we systematically examine the alignment of the distributions of adversarial examples when subjected to the process of transformation using diffusion models. The efficacy of this approach is assessed across CIFAR-10 and ImageNet datasets, including varying image sizes in the latter. The results demonstrate a notable capacity to discriminate effectively between benign and attacked images, providing compelling evidence that adversarial instances do not align with the learned manifold of the DMs.
title Adversarial Examples are Misaligned in Diffusion Model Manifolds
topic Computer Vision and Pattern Recognition
Cryptography and Security
url https://arxiv.org/abs/2401.06637