SD-NAE: Generating Natural Adversarial Examples with Stable Diffusion

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lin, Yueqian, Zhang, Jingyang, Chen, Yiran, Li, Hai
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911876376428544
author Lin, Yueqian
Zhang, Jingyang
Chen, Yiran
Li, Hai
author_facet Lin, Yueqian
Zhang, Jingyang
Chen, Yiran
Li, Hai
contents Natural Adversarial Examples (NAEs), images arising naturally from the environment and capable of deceiving classifiers, are instrumental in robustly evaluating and identifying vulnerabilities in trained models. In this work, unlike prior works that passively collect NAEs from real images, we propose to actively synthesize NAEs using the state-of-the-art Stable Diffusion. Specifically, our method formulates a controlled optimization process, where we perturb the token embedding that corresponds to a specified class to generate NAEs. This generation process is guided by the gradient of loss from the target classifier, ensuring that the created image closely mimics the ground-truth class yet fools the classifier. Named SD-NAE (Stable Diffusion for Natural Adversarial Examples), our innovative method is effective in producing valid and useful NAEs, which is demonstrated through a meticulously designed experiment. Code is available at https://github.com/linyueqian/SD-NAE.
format Preprint
id arxiv_https___arxiv_org_abs_2311_12981
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SD-NAE: Generating Natural Adversarial Examples with Stable Diffusion
Lin, Yueqian
Zhang, Jingyang
Chen, Yiran
Li, Hai
Computer Vision and Pattern Recognition
Natural Adversarial Examples (NAEs), images arising naturally from the environment and capable of deceiving classifiers, are instrumental in robustly evaluating and identifying vulnerabilities in trained models. In this work, unlike prior works that passively collect NAEs from real images, we propose to actively synthesize NAEs using the state-of-the-art Stable Diffusion. Specifically, our method formulates a controlled optimization process, where we perturb the token embedding that corresponds to a specified class to generate NAEs. This generation process is guided by the gradient of loss from the target classifier, ensuring that the created image closely mimics the ground-truth class yet fools the classifier. Named SD-NAE (Stable Diffusion for Natural Adversarial Examples), our innovative method is effective in producing valid and useful NAEs, which is demonstrated through a meticulously designed experiment. Code is available at https://github.com/linyueqian/SD-NAE.
title SD-NAE: Generating Natural Adversarial Examples with Stable Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.12981