GuardT2I: Defending Text-to-Image Models from Adversarial Prompts

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Yijun, Gao, Ruiyuan, Yang, Xiao, Zhong, Jianyuan, Xu, Qiang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910678010298368
author Yang, Yijun
Gao, Ruiyuan
Yang, Xiao
Zhong, Jianyuan
Xu, Qiang
author_facet Yang, Yijun
Gao, Ruiyuan
Yang, Xiao
Zhong, Jianyuan
Xu, Qiang
contents Recent advancements in Text-to-Image (T2I) models have raised significant safety concerns about their potential misuse for generating inappropriate or Not-Safe-For-Work (NSFW) contents, despite existing countermeasures such as NSFW classifiers or model fine-tuning for inappropriate concept removal. Addressing this challenge, our study unveils GuardT2I, a novel moderation framework that adopts a generative approach to enhance T2I models' robustness against adversarial prompts. Instead of making a binary classification, GuardT2I utilizes a Large Language Model (LLM) to conditionally transform text guidance embeddings within the T2I models into natural language for effective adversarial prompt detection, without compromising the models' inherent performance. Our extensive experiments reveal that GuardT2I outperforms leading commercial solutions like OpenAI-Moderation and Microsoft Azure Moderator by a significant margin across diverse adversarial scenarios. Our framework is available at https://github.com/cure-lab/GuardT2I.
format Preprint
id arxiv_https___arxiv_org_abs_2403_01446
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GuardT2I: Defending Text-to-Image Models from Adversarial Prompts
Yang, Yijun
Gao, Ruiyuan
Yang, Xiao
Zhong, Jianyuan
Xu, Qiang
Computer Vision and Pattern Recognition
Recent advancements in Text-to-Image (T2I) models have raised significant safety concerns about their potential misuse for generating inappropriate or Not-Safe-For-Work (NSFW) contents, despite existing countermeasures such as NSFW classifiers or model fine-tuning for inappropriate concept removal. Addressing this challenge, our study unveils GuardT2I, a novel moderation framework that adopts a generative approach to enhance T2I models' robustness against adversarial prompts. Instead of making a binary classification, GuardT2I utilizes a Large Language Model (LLM) to conditionally transform text guidance embeddings within the T2I models into natural language for effective adversarial prompt detection, without compromising the models' inherent performance. Our extensive experiments reveal that GuardT2I outperforms leading commercial solutions like OpenAI-Moderation and Microsoft Azure Moderator by a significant margin across diverse adversarial scenarios. Our framework is available at https://github.com/cure-lab/GuardT2I.
title GuardT2I: Defending Text-to-Image Models from Adversarial Prompts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.01446