A Unified Framework for LLM Watermarks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gloaguen, Thibaud, Staab, Robin, Jovanović, Nikola, Vechev, Martin
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918325908406272
author Gloaguen, Thibaud
Staab, Robin
Jovanović, Nikola
Vechev, Martin
author_facet Gloaguen, Thibaud
Staab, Robin
Jovanović, Nikola
Vechev, Martin
contents LLM watermarks allow tracing AI-generated texts by inserting a detectable signal into their generated content. Recent works have proposed a wide range of watermarking algorithms, each with distinct designs, usually built using a bottom-up approach. Crucially, there is no general and principled formulation for LLM watermarking. In this work, we show that most existing and widely used watermarking schemes can in fact be derived from a principled constrained optimization problem. Our formulation unifies existing watermarking methods and explicitly reveals the constraints that each method optimizes. In particular, it highlights an understudied quality-diversity-power trade-off. At the same time, our framework also provides a principled approach for designing novel watermarking schemes tailored to specific requirements. For instance, it allows us to directly use perplexity as a proxy for quality, and derive new schemes that are optimal with respect to this constraint. Our experimental evaluation validates our framework: watermarking schemes derived from a given constraint consistently maximize detection power with respect to that constraint.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06754
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Unified Framework for LLM Watermarks
Gloaguen, Thibaud
Staab, Robin
Jovanović, Nikola
Vechev, Martin
Cryptography and Security
Artificial Intelligence
Machine Learning
LLM watermarks allow tracing AI-generated texts by inserting a detectable signal into their generated content. Recent works have proposed a wide range of watermarking algorithms, each with distinct designs, usually built using a bottom-up approach. Crucially, there is no general and principled formulation for LLM watermarking. In this work, we show that most existing and widely used watermarking schemes can in fact be derived from a principled constrained optimization problem. Our formulation unifies existing watermarking methods and explicitly reveals the constraints that each method optimizes. In particular, it highlights an understudied quality-diversity-power trade-off. At the same time, our framework also provides a principled approach for designing novel watermarking schemes tailored to specific requirements. For instance, it allows us to directly use perplexity as a proxy for quality, and derive new schemes that are optimal with respect to this constraint. Our experimental evaluation validates our framework: watermarking schemes derived from a given constraint consistently maximize detection power with respect to that constraint.
title A Unified Framework for LLM Watermarks
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2602.06754