Watermarking Diffusion Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gloaguen, Thibaud, Staab, Robin, Jovanović, Nikola, Vechev, Martin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912912496394240
author Gloaguen, Thibaud
Staab, Robin
Jovanović, Nikola
Vechev, Martin
author_facet Gloaguen, Thibaud
Staab, Robin
Jovanović, Nikola
Vechev, Martin
contents We introduce the first watermark tailored for diffusion language models (DLMs), an emergent LLM paradigm able to generate tokens in arbitrary order, in contrast to standard autoregressive language models (ARLMs) which generate tokens sequentially. While there has been much work in ARLM watermarking, a key challenge when attempting to apply these schemes directly to the DLM setting is that they rely on previously generated tokens, which are not always available with DLM generation. In this work we address this challenge by: (i) applying the watermark in expectation over the context even when some context tokens are yet to be determined, and (ii) promoting tokens which increase the watermark strength when used as context for other tokens. This is accomplished while keeping the watermark detector unchanged. Our experimental evaluation demonstrates that the DLM watermark leads to a >99% true positive rate with minimal quality impact and achieves similar robustness to existing ARLM watermarks, enabling for the first time reliable DLM watermarking.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24368
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Watermarking Diffusion Language Models
Gloaguen, Thibaud
Staab, Robin
Jovanović, Nikola
Vechev, Martin
Machine Learning
Artificial Intelligence
Cryptography and Security
We introduce the first watermark tailored for diffusion language models (DLMs), an emergent LLM paradigm able to generate tokens in arbitrary order, in contrast to standard autoregressive language models (ARLMs) which generate tokens sequentially. While there has been much work in ARLM watermarking, a key challenge when attempting to apply these schemes directly to the DLM setting is that they rely on previously generated tokens, which are not always available with DLM generation. In this work we address this challenge by: (i) applying the watermark in expectation over the context even when some context tokens are yet to be determined, and (ii) promoting tokens which increase the watermark strength when used as context for other tokens. This is accomplished while keeping the watermark detector unchanged. Our experimental evaluation demonstrates that the DLM watermark leads to a >99% true positive rate with minimal quality impact and achieves similar robustness to existing ARLM watermarks, enabling for the first time reliable DLM watermarking.
title Watermarking Diffusion Language Models
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2509.24368