JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zheng, Haolun, He, Yu, Chen, Tailun, Shao, Shuo, Chu, Zhixuan, Zhou, Hongbin, Tao, Lan, Qin, Zhan, Ren, Kui
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915891511296000
author Zheng, Haolun
He, Yu
Chen, Tailun
Shao, Shuo
Chu, Zhixuan
Zhou, Hongbin
Tao, Lan
Qin, Zhan
Ren, Kui
author_facet Zheng, Haolun
He, Yu
Chen, Tailun
Shao, Shuo
Chu, Zhixuan
Zhou, Hongbin
Tao, Lan
Qin, Zhan
Ren, Kui
contents Text-to-image (T2I) models such as Stable Diffusion and DALLE remain susceptible to generating harmful or Not-Safe-For-Work (NSFW) content under jailbreak attacks despite deployed safety filters. Existing jailbreak attacks either rely on proxy-loss optimization instead of the true end-to-end objective, or depend on large-scale and costly RL-trained generators. Motivated by these limitations, we propose JANUS , a lightweight framework that formulates jailbreak as optimizing a structured prompt distribution under a black-box, end-to-end reward from the T2I system and its safety filters. JANUS replaces a high-capacity generator with a low-dimensional mixing policy over two semantically anchored prompt distributions, enabling efficient exploration while preserving the target semantics. On modern T2I models, we outperform state-of-the-art jailbreak methods, improving ASR-8 from 25.30% to 43.15% on Stable Diffusion 3.5 Large Turbo with consistently higher CLIP and NSFW scores. JANUS succeeds across both open-source and commercial models. These findings expose structural weaknesses in current T2I safety pipelines and motivate stronger, distribution-aware defenses. Warning: This paper contains model outputs that may be offensive.
format Preprint
id arxiv_https___arxiv_org_abs_2603_21208
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization
Zheng, Haolun
He, Yu
Chen, Tailun
Shao, Shuo
Chu, Zhixuan
Zhou, Hongbin
Tao, Lan
Qin, Zhan
Ren, Kui
Computer Vision and Pattern Recognition
Machine Learning
Text-to-image (T2I) models such as Stable Diffusion and DALLE remain susceptible to generating harmful or Not-Safe-For-Work (NSFW) content under jailbreak attacks despite deployed safety filters. Existing jailbreak attacks either rely on proxy-loss optimization instead of the true end-to-end objective, or depend on large-scale and costly RL-trained generators. Motivated by these limitations, we propose JANUS , a lightweight framework that formulates jailbreak as optimizing a structured prompt distribution under a black-box, end-to-end reward from the T2I system and its safety filters. JANUS replaces a high-capacity generator with a low-dimensional mixing policy over two semantically anchored prompt distributions, enabling efficient exploration while preserving the target semantics. On modern T2I models, we outperform state-of-the-art jailbreak methods, improving ASR-8 from 25.30% to 43.15% on Stable Diffusion 3.5 Large Turbo with consistently higher CLIP and NSFW scores. JANUS succeeds across both open-source and commercial models. These findings expose structural weaknesses in current T2I safety pipelines and motivate stronger, distribution-aware defenses. Warning: This paper contains model outputs that may be offensive.
title JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2603.21208