AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Yaopei, Cao, Yuanpu, Cao, Bochuan, Chang, Yurui, Chen, Jinghui, Lin, Lu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915490680537088
author Zeng, Yaopei
Cao, Yuanpu
Cao, Bochuan
Chang, Yurui
Chen, Jinghui
Lin, Lu
author_facet Zeng, Yaopei
Cao, Yuanpu
Cao, Bochuan
Chang, Yurui
Chen, Jinghui
Lin, Lu
contents Recent advances in diffusion models have significantly enhanced the quality of image synthesis, yet they have also introduced serious safety concerns, particularly the generation of Not Safe for Work (NSFW) content. Previous research has demonstrated that adversarial prompts can be used to generate NSFW content. However, such adversarial text prompts are often easily detectable by text-based filters, limiting their efficacy. In this paper, we expose a previously overlooked vulnerability: adversarial image attacks targeting Image-to-Image (I2I) diffusion models. We propose AdvI2I, a novel framework that manipulates input images to induce diffusion models to generate NSFW content. By optimizing a generator to craft adversarial images, AdvI2I circumvents existing defense mechanisms, such as Safe Latent Diffusion (SLD), without altering the text prompts. Furthermore, we introduce AdvI2I-Adaptive, an enhanced version that adapts to potential countermeasures and minimizes the resemblance between adversarial images and NSFW concept embeddings, making the attack more resilient against defenses. Through extensive experiments, we demonstrate that both AdvI2I and AdvI2I-Adaptive can effectively bypass current safeguards, highlighting the urgent need for stronger security measures to address the misuse of I2I diffusion models.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21471
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion models
Zeng, Yaopei
Cao, Yuanpu
Cao, Bochuan
Chang, Yurui
Chen, Jinghui
Lin, Lu
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advances in diffusion models have significantly enhanced the quality of image synthesis, yet they have also introduced serious safety concerns, particularly the generation of Not Safe for Work (NSFW) content. Previous research has demonstrated that adversarial prompts can be used to generate NSFW content. However, such adversarial text prompts are often easily detectable by text-based filters, limiting their efficacy. In this paper, we expose a previously overlooked vulnerability: adversarial image attacks targeting Image-to-Image (I2I) diffusion models. We propose AdvI2I, a novel framework that manipulates input images to induce diffusion models to generate NSFW content. By optimizing a generator to craft adversarial images, AdvI2I circumvents existing defense mechanisms, such as Safe Latent Diffusion (SLD), without altering the text prompts. Furthermore, we introduce AdvI2I-Adaptive, an enhanced version that adapts to potential countermeasures and minimizes the resemblance between adversarial images and NSFW concept embeddings, making the attack more resilient against defenses. Through extensive experiments, we demonstrate that both AdvI2I and AdvI2I-Adaptive can effectively bypass current safeguards, highlighting the urgent need for stronger security measures to address the misuse of I2I diffusion models.
title AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2410.21471