Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shan, Shawn, Ding, Wenxin, Passananti, Josephine, Wu, Stanley, Zheng, Haitao, Zhao, Ben Y.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911857721212928
author Shan, Shawn
Ding, Wenxin
Passananti, Josephine
Wu, Stanley
Zheng, Haitao
Zhao, Ben Y.
author_facet Shan, Shawn
Ding, Wenxin
Passananti, Josephine
Wu, Stanley
Zheng, Haitao
Zhao, Ben Y.
contents Data poisoning attacks manipulate training data to introduce unexpected behaviors into machine learning models at training time. For text-to-image generative models with massive training datasets, current understanding of poisoning attacks suggests that a successful attack would require injecting millions of poison samples into their training pipeline. In this paper, we show that poisoning attacks can be successful on generative models. We observe that training data per concept can be quite limited in these models, making them vulnerable to prompt-specific poisoning attacks, which target a model's ability to respond to individual prompts. We introduce Nightshade, an optimized prompt-specific poisoning attack where poison samples look visually identical to benign images with matching text prompts. Nightshade poison samples are also optimized for potency and can corrupt an Stable Diffusion SDXL prompt in <100 poison samples. Nightshade poison effects "bleed through" to related concepts, and multiple attacks can composed together in a single prompt. Surprisingly, we show that a moderate number of Nightshade attacks can destabilize general features in a text-to-image generative model, effectively disabling its ability to generate meaningful images. Finally, we propose the use of Nightshade and similar tools as a last defense for content creators against web scrapers that ignore opt-out/do-not-crawl directives, and discuss possible implications for model trainers and content creators.
format Preprint
id arxiv_https___arxiv_org_abs_2310_13828
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models
Shan, Shawn
Ding, Wenxin
Passananti, Josephine
Wu, Stanley
Zheng, Haitao
Zhao, Ben Y.
Cryptography and Security
Artificial Intelligence
Data poisoning attacks manipulate training data to introduce unexpected behaviors into machine learning models at training time. For text-to-image generative models with massive training datasets, current understanding of poisoning attacks suggests that a successful attack would require injecting millions of poison samples into their training pipeline. In this paper, we show that poisoning attacks can be successful on generative models. We observe that training data per concept can be quite limited in these models, making them vulnerable to prompt-specific poisoning attacks, which target a model's ability to respond to individual prompts. We introduce Nightshade, an optimized prompt-specific poisoning attack where poison samples look visually identical to benign images with matching text prompts. Nightshade poison samples are also optimized for potency and can corrupt an Stable Diffusion SDXL prompt in <100 poison samples. Nightshade poison effects "bleed through" to related concepts, and multiple attacks can composed together in a single prompt. Surprisingly, we show that a moderate number of Nightshade attacks can destabilize general features in a text-to-image generative model, effectively disabling its ability to generate meaningful images. Finally, we propose the use of Nightshade and similar tools as a last defense for content creators against web scrapers that ignore opt-out/do-not-crawl directives, and discuss possible implications for model trainers and content creators.
title Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2310.13828