Watermarking Should Be Treated as a Monitoring Primitive

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aremu, Toluwani, Lukas, Nils, Zhang, Jie
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910220979011584
author Aremu, Toluwani
Lukas, Nils
Zhang, Jie
author_facet Aremu, Toluwani
Lukas, Nils
Zhang, Jie
contents Watermarking is widely proposed for provenance, attribution, and safety monitoring in generative models, yet is typically evaluated only under adversaries who attempt to evade detection or induce false positives at the level of individual samples. We argue that watermarking should be treated as a monitoring primitive, and that internal monitoring is unavoidable given per-entity attribution keys and messages, as well as detector access. We introduce an observer-based threat model in which observers can aggregate watermark signals across outputs to infer entity-level information, showing that even zero-bit watermarking enables attribution under multi-key settings. We further show that external monitoring can emerge over time from persistent, key-dependent statistical structure, although this depends on watermark design and may be mitigated by distribution-preserving or undetectable schemes. Our findings reveal a fundamental dual-use tension between attribution and monitoring, motivating evaluation of watermarking beyond per-sample robustness to account for aggregation and observer-based capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13095
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Watermarking Should Be Treated as a Monitoring Primitive
Aremu, Toluwani
Lukas, Nils
Zhang, Jie
Cryptography and Security
Artificial Intelligence
Computers and Society
Machine Learning
Watermarking is widely proposed for provenance, attribution, and safety monitoring in generative models, yet is typically evaluated only under adversaries who attempt to evade detection or induce false positives at the level of individual samples. We argue that watermarking should be treated as a monitoring primitive, and that internal monitoring is unavoidable given per-entity attribution keys and messages, as well as detector access. We introduce an observer-based threat model in which observers can aggregate watermark signals across outputs to infer entity-level information, showing that even zero-bit watermarking enables attribution under multi-key settings. We further show that external monitoring can emerge over time from persistent, key-dependent statistical structure, although this depends on watermark design and may be mitigated by distribution-preserving or undetectable schemes. Our findings reveal a fundamental dual-use tension between attribution and monitoring, motivating evaluation of watermarking beyond per-sample robustness to account for aggregation and observer-based capabilities.
title Watermarking Should Be Treated as a Monitoring Primitive
topic Cryptography and Security
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2605.13095