The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wen, Zichen, Qu, Jiashu, Chen, Zhaorun, Lu, Xiaoya, Liu, Dongrui, Liu, Zhiyuan, Wu, Ruixi, Yang, Yicun, Jin, Xiangqi, Xu, Haoyun, Liu, Xuyang, Li, Weijia, Lu, Chaochao, Shao, Jing, He, Conghui, Zhang, Linfeng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918329533333504
author Wen, Zichen
Qu, Jiashu
Chen, Zhaorun
Lu, Xiaoya
Liu, Dongrui
Liu, Zhiyuan
Wu, Ruixi
Yang, Yicun
Jin, Xiangqi
Xu, Haoyun
Liu, Xuyang
Li, Weijia
Lu, Chaochao
Shao, Jing
He, Conghui
Zhang, Linfeng
author_facet Wen, Zichen
Qu, Jiashu
Chen, Zhaorun
Lu, Xiaoya
Liu, Dongrui
Liu, Zhiyuan
Wu, Ruixi
Yang, Yicun
Jin, Xiangqi
Xu, Haoyun
Liu, Xuyang
Li, Weijia
Lu, Chaochao
Shao, Jing
He, Conghui
Zhang, Linfeng
contents Diffusion-based large language models (dLLMs) have recently emerged as a powerful alternative to autoregressive LLMs, offering faster inference and greater interactivity via parallel decoding and bidirectional modeling. However, despite strong performance in code generation and text infilling, we identify a fundamental safety concern: existing alignment mechanisms fail to safeguard dLLMs against context-aware, masked-input adversarial prompts, exposing novel vulnerabilities. To this end, we present DIJA, the first systematic study and jailbreak attack framework that exploits unique safety weaknesses of dLLMs. Specifically, our proposed DIJA constructs adversarial interleaved mask-text prompts that exploit the text generation mechanisms of dLLMs, i.e., bidirectional modeling and parallel decoding. Bidirectional modeling drives the model to produce contextually consistent outputs for masked spans, even when harmful, while parallel decoding limits model dynamic filtering and rejection sampling of unsafe content. This causes standard alignment mechanisms to fail, enabling harmful completions in alignment-tuned dLLMs, even when harmful behaviors or unsafe instructions are directly exposed in the prompt. Through comprehensive experiments, we demonstrate that DIJA significantly outperforms existing jailbreak methods, exposing a previously overlooked threat surface in dLLM architectures. Notably, our method achieves up to 100% keyword-based ASR on Dream-Instruct, surpassing the strongest prior baseline, ReNeLLM, by up to 78.5% in evaluator-based ASR on JailbreakBench and by 37.7 points in StrongREJECT score, while requiring no rewriting or hiding of harmful content in the jailbreak prompt. Our findings underscore the urgent need for rethinking safety alignment in this emerging class of language models. Code is available at https://github.com/ZichenWen1/DIJA.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11097
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
Wen, Zichen
Qu, Jiashu
Chen, Zhaorun
Lu, Xiaoya
Liu, Dongrui
Liu, Zhiyuan
Wu, Ruixi
Yang, Yicun
Jin, Xiangqi
Xu, Haoyun
Liu, Xuyang
Li, Weijia
Lu, Chaochao
Shao, Jing
He, Conghui
Zhang, Linfeng
Computation and Language
Diffusion-based large language models (dLLMs) have recently emerged as a powerful alternative to autoregressive LLMs, offering faster inference and greater interactivity via parallel decoding and bidirectional modeling. However, despite strong performance in code generation and text infilling, we identify a fundamental safety concern: existing alignment mechanisms fail to safeguard dLLMs against context-aware, masked-input adversarial prompts, exposing novel vulnerabilities. To this end, we present DIJA, the first systematic study and jailbreak attack framework that exploits unique safety weaknesses of dLLMs. Specifically, our proposed DIJA constructs adversarial interleaved mask-text prompts that exploit the text generation mechanisms of dLLMs, i.e., bidirectional modeling and parallel decoding. Bidirectional modeling drives the model to produce contextually consistent outputs for masked spans, even when harmful, while parallel decoding limits model dynamic filtering and rejection sampling of unsafe content. This causes standard alignment mechanisms to fail, enabling harmful completions in alignment-tuned dLLMs, even when harmful behaviors or unsafe instructions are directly exposed in the prompt. Through comprehensive experiments, we demonstrate that DIJA significantly outperforms existing jailbreak methods, exposing a previously overlooked threat surface in dLLM architectures. Notably, our method achieves up to 100% keyword-based ASR on Dream-Instruct, surpassing the strongest prior baseline, ReNeLLM, by up to 78.5% in evaluator-based ASR on JailbreakBench and by 37.7 points in StrongREJECT score, while requiring no rewriting or hiding of harmful content in the jailbreak prompt. Our findings underscore the urgent need for rethinking safety alignment in this emerging class of language models. Code is available at https://github.com/ZichenWen1/DIJA.
title The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
topic Computation and Language
url https://arxiv.org/abs/2507.11097