AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yiming, Chen, Jiahao, Li, Qingming, Zhang, Tong, Zeng, Rui, Yang, Xing, Ji, Shouling
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917133220315136
author Wang, Yiming
Chen, Jiahao
Li, Qingming
Zhang, Tong
Zeng, Rui
Yang, Xing
Ji, Shouling
author_facet Wang, Yiming
Chen, Jiahao
Li, Qingming
Zhang, Tong
Zeng, Rui
Yang, Xing
Ji, Shouling
contents As text-to-image (T2I) models advance and gain widespread adoption, their associated safety concerns are becoming increasingly critical. Malicious users exploit these models to generate Not-Safe-for-Work (NSFW) images using harmful or adversarial prompts, underscoring the need for effective safeguards to ensure the integrity and compliance of model outputs. However, existing detection methods often exhibit low accuracy and inefficiency. In this paper, we propose AEIOU, a defense framework that is adaptable, efficient, interpretable, optimizable, and unified against NSFW prompts in T2I models. AEIOU extracts NSFW features from the hidden states of the model's text encoder, utilizing the separable nature of these features to detect NSFW prompts. The detection process is efficient, requiring minimal inference time. AEIOU also offers real-time interpretation of results and supports optimization through data augmentation techniques. The framework is versatile, accommodating various T2I architectures. Our extensive experiments show that AEIOU significantly outperforms both commercial and open-source moderation tools, achieving over 95\% accuracy across all datasets and improving efficiency by at least tenfold. It effectively counters adaptive attacks and excels in few-shot and multi-label scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18123
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models
Wang, Yiming
Chen, Jiahao
Li, Qingming
Zhang, Tong
Zeng, Rui
Yang, Xing
Ji, Shouling
Cryptography and Security
Computation and Language
As text-to-image (T2I) models advance and gain widespread adoption, their associated safety concerns are becoming increasingly critical. Malicious users exploit these models to generate Not-Safe-for-Work (NSFW) images using harmful or adversarial prompts, underscoring the need for effective safeguards to ensure the integrity and compliance of model outputs. However, existing detection methods often exhibit low accuracy and inefficiency. In this paper, we propose AEIOU, a defense framework that is adaptable, efficient, interpretable, optimizable, and unified against NSFW prompts in T2I models. AEIOU extracts NSFW features from the hidden states of the model's text encoder, utilizing the separable nature of these features to detect NSFW prompts. The detection process is efficient, requiring minimal inference time. AEIOU also offers real-time interpretation of results and supports optimization through data augmentation techniques. The framework is versatile, accommodating various T2I architectures. Our extensive experiments show that AEIOU significantly outperforms both commercial and open-source moderation tools, achieving over 95\% accuracy across all datasets and improving efficiency by at least tenfold. It effectively counters adaptive attacks and excels in few-shot and multi-label scenarios.
title AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2412.18123