A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jeung, Wonje, Yoon, Sangyeon, Cho, Yoonjun, Jeon, Dongjae, Shin, Sangwoo, Hong, Hyesoo, No, Albert
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910009340723200
author Jeung, Wonje
Yoon, Sangyeon
Cho, Yoonjun
Jeon, Dongjae
Shin, Sangwoo
Hong, Hyesoo
No, Albert
author_facet Jeung, Wonje
Yoon, Sangyeon
Cho, Yoonjun
Jeon, Dongjae
Shin, Sangwoo
Hong, Hyesoo
No, Albert
contents Diffusion large language models (dLLMs) enable any-order generation, but this flexibility enlarges the attack surface: harmful spans may appear at arbitrary positions, and template-based prefilling attacks such as DIJA bypass response-level refusals. We introduce A2D (Any-Order, Any-Step Defense), a token-level alignment method that aligns dLLMs to emit an [EOS] refusal signal whenever harmful content arises. By aligning safety directly at the token-level under randomized masking, A2D achieves robustness to both any-decoding-order and any-step prefilling attacks under various conditions. It also enables real-time monitoring: dLLMs may begin a response but automatically terminate if unsafe continuation emerges. On safety benchmarks, A2D consistently prevents the generation of harmful outputs, slashing DIJA success rates from over 80% to near-zero (1.3% on LLaDA-8B-Instruct, 0.0% on Dream-v0-Instruct-7B), and thresholded [EOS] probabilities allow early rejection, yielding up to 19.3x faster safe termination.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23286
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
Jeung, Wonje
Yoon, Sangyeon
Cho, Yoonjun
Jeon, Dongjae
Shin, Sangwoo
Hong, Hyesoo
No, Albert
Computation and Language
Artificial Intelligence
Diffusion large language models (dLLMs) enable any-order generation, but this flexibility enlarges the attack surface: harmful spans may appear at arbitrary positions, and template-based prefilling attacks such as DIJA bypass response-level refusals. We introduce A2D (Any-Order, Any-Step Defense), a token-level alignment method that aligns dLLMs to emit an [EOS] refusal signal whenever harmful content arises. By aligning safety directly at the token-level under randomized masking, A2D achieves robustness to both any-decoding-order and any-step prefilling attacks under various conditions. It also enables real-time monitoring: dLLMs may begin a response but automatically terminate if unsafe continuation emerges. On safety benchmarks, A2D consistently prevents the generation of harmful outputs, slashing DIJA success rates from over 80% to near-zero (1.3% on LLaDA-8B-Instruct, 0.0% on Dream-v0-Instruct-7B), and thresholded [EOS] probabilities allow early rejection, yielding up to 19.3x faster safe termination.
title A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.23286