JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915891511296000 |
|---|---|
| author | Zheng, Haolun He, Yu Chen, Tailun Shao, Shuo Chu, Zhixuan Zhou, Hongbin Tao, Lan Qin, Zhan Ren, Kui |
| author_facet | Zheng, Haolun He, Yu Chen, Tailun Shao, Shuo Chu, Zhixuan Zhou, Hongbin Tao, Lan Qin, Zhan Ren, Kui |
| contents | Text-to-image (T2I) models such as Stable Diffusion and DALLE remain susceptible to generating harmful or Not-Safe-For-Work (NSFW) content under jailbreak attacks despite deployed safety filters. Existing jailbreak attacks either rely on proxy-loss optimization instead of the true end-to-end objective, or depend on large-scale and costly RL-trained generators. Motivated by these limitations, we propose JANUS , a lightweight framework that formulates jailbreak as optimizing a structured prompt distribution under a black-box, end-to-end reward from the T2I system and its safety filters. JANUS replaces a high-capacity generator with a low-dimensional mixing policy over two semantically anchored prompt distributions, enabling efficient exploration while preserving the target semantics. On modern T2I models, we outperform state-of-the-art jailbreak methods, improving ASR-8 from 25.30% to 43.15% on Stable Diffusion 3.5 Large Turbo with consistently higher CLIP and NSFW scores. JANUS succeeds across both open-source and commercial models. These findings expose structural weaknesses in current T2I safety pipelines and motivate stronger, distribution-aware defenses. Warning: This paper contains model outputs that may be offensive. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_21208 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization Zheng, Haolun He, Yu Chen, Tailun Shao, Shuo Chu, Zhixuan Zhou, Hongbin Tao, Lan Qin, Zhan Ren, Kui Computer Vision and Pattern Recognition Machine Learning Text-to-image (T2I) models such as Stable Diffusion and DALLE remain susceptible to generating harmful or Not-Safe-For-Work (NSFW) content under jailbreak attacks despite deployed safety filters. Existing jailbreak attacks either rely on proxy-loss optimization instead of the true end-to-end objective, or depend on large-scale and costly RL-trained generators. Motivated by these limitations, we propose JANUS , a lightweight framework that formulates jailbreak as optimizing a structured prompt distribution under a black-box, end-to-end reward from the T2I system and its safety filters. JANUS replaces a high-capacity generator with a low-dimensional mixing policy over two semantically anchored prompt distributions, enabling efficient exploration while preserving the target semantics. On modern T2I models, we outperform state-of-the-art jailbreak methods, improving ASR-8 from 25.30% to 43.15% on Stable Diffusion 3.5 Large Turbo with consistently higher CLIP and NSFW scores. JANUS succeeds across both open-source and commercial models. These findings expose structural weaknesses in current T2I safety pipelines and motivate stronger, distribution-aware defenses. Warning: This paper contains model outputs that may be offensive. |
| title | JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2603.21208 |