Jailbreaking Safeguarded Text-to-Image Models via Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jiang, Zhengyuan, Hu, Yuepeng, Yang, Yuchen, Cao, Yinzhi, Gong, Neil Zhenqiang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909983919046656
author Jiang, Zhengyuan
Hu, Yuepeng
Yang, Yuchen
Cao, Yinzhi
Gong, Neil Zhenqiang
author_facet Jiang, Zhengyuan
Hu, Yuepeng
Yang, Yuchen
Cao, Yinzhi
Gong, Neil Zhenqiang
contents Text-to-Image models may generate harmful content, such as pornographic images, particularly when unsafe prompts are submitted. To address this issue, safety filters are often added on top of text-to-image models, or the models themselves are aligned to reduce harmful outputs. However, these defenses remain vulnerable when an attacker strategically designs adversarial prompts to bypass these safety guardrails. In this work, we propose \alg, a method to jailbreak text-to-image models with safety guardrails using a fine-tuned large language model. Unlike other query-based jailbreak attacks that require repeated queries to the target model, our attack generates adversarial prompts efficiently after fine-tuning our AttackLLM. We evaluate our method on three datasets of unsafe prompts and against five safety guardrails. Our results demonstrate that our approach effectively bypasses safety guardrails, outperforms existing no-box attacks, and also facilitates other query-based attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_01839
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Jailbreaking Safeguarded Text-to-Image Models via Large Language Models
Jiang, Zhengyuan
Hu, Yuepeng
Yang, Yuchen
Cao, Yinzhi
Gong, Neil Zhenqiang
Cryptography and Security
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Text-to-Image models may generate harmful content, such as pornographic images, particularly when unsafe prompts are submitted. To address this issue, safety filters are often added on top of text-to-image models, or the models themselves are aligned to reduce harmful outputs. However, these defenses remain vulnerable when an attacker strategically designs adversarial prompts to bypass these safety guardrails. In this work, we propose \alg, a method to jailbreak text-to-image models with safety guardrails using a fine-tuned large language model. Unlike other query-based jailbreak attacks that require repeated queries to the target model, our attack generates adversarial prompts efficiently after fine-tuning our AttackLLM. We evaluate our method on three datasets of unsafe prompts and against five safety guardrails. Our results demonstrate that our approach effectively bypasses safety guardrails, outperforms existing no-box attacks, and also facilitates other query-based attacks.
title Jailbreaking Safeguarded Text-to-Image Models via Large Language Models
topic Cryptography and Security
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.01839