MMA-Diffusion: MultiModal Attack on Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yijun, Gao, Ruiyuan, Wang, Xiaosen, Ho, Tsung-Yi, Xu, Nan, Xu, Qiang
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914735310503936
author Yang, Yijun
Gao, Ruiyuan
Wang, Xiaosen
Ho, Tsung-Yi
Xu, Nan
Xu, Qiang
author_facet Yang, Yijun
Gao, Ruiyuan
Wang, Xiaosen
Ho, Tsung-Yi
Xu, Nan
Xu, Qiang
contents In recent years, Text-to-Image (T2I) models have seen remarkable advancements, gaining widespread adoption. However, this progress has inadvertently opened avenues for potential misuse, particularly in generating inappropriate or Not-Safe-For-Work (NSFW) content. Our work introduces MMA-Diffusion, a framework that presents a significant and realistic threat to the security of T2I models by effectively circumventing current defensive measures in both open-source models and commercial online services. Unlike previous approaches, MMA-Diffusion leverages both textual and visual modalities to bypass safeguards like prompt filters and post-hoc safety checkers, thus exposing and highlighting the vulnerabilities in existing defense mechanisms.
format Preprint
id arxiv_https___arxiv_org_abs_2311_17516
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MMA-Diffusion: MultiModal Attack on Diffusion Models
Yang, Yijun
Gao, Ruiyuan
Wang, Xiaosen
Ho, Tsung-Yi
Xu, Nan
Xu, Qiang
Cryptography and Security
Computer Vision and Pattern Recognition
In recent years, Text-to-Image (T2I) models have seen remarkable advancements, gaining widespread adoption. However, this progress has inadvertently opened avenues for potential misuse, particularly in generating inappropriate or Not-Safe-For-Work (NSFW) content. Our work introduces MMA-Diffusion, a framework that presents a significant and realistic threat to the security of T2I models by effectively circumventing current defensive measures in both open-source models and commercial online services. Unlike previous approaches, MMA-Diffusion leverages both textual and visual modalities to bypass safeguards like prompt filters and post-hoc safety checkers, thus exposing and highlighting the vulnerabilities in existing defense mechanisms.
title MMA-Diffusion: MultiModal Attack on Diffusion Models
topic Cryptography and Security
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.17516