Sandcastles in the Storm: Revisiting the (Im)possibility of Strong Watermarking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Harel-Canada, Fabrice Y, Erol, Boran, Choi, Connor, Liu, Jason, Song, Gary Jiarui, Peng, Nanyun, Sahai, Amit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908359742980096
author Harel-Canada, Fabrice Y
Erol, Boran
Choi, Connor
Liu, Jason
Song, Gary Jiarui
Peng, Nanyun
Sahai, Amit
author_facet Harel-Canada, Fabrice Y
Erol, Boran
Choi, Connor
Liu, Jason
Song, Gary Jiarui
Peng, Nanyun
Sahai, Amit
contents Watermarking AI-generated text is critical for combating misuse. Yet recent theoretical work argues that any watermark can be erased via random walk attacks that perturb text while preserving quality. However, such attacks rely on two key assumptions: (1) rapid mixing (watermarks dissolve quickly under perturbations) and (2) reliable quality preservation (automated quality oracles perfectly guide edits). Through large-scale experiments and human-validated assessments, we find mixing is slow: 100% of perturbed texts retain traces of their origin after hundreds of edits, defying rapid mixing. Oracles falter, as state-of-the-art quality detectors misjudge edits (77% accuracy), compounding errors during attacks. Ultimately, attacks underperform: automated walks remove watermarks just 26% of the time -- dropping to 10% under human quality review. These findings challenge the inevitability of watermark removal. Instead, practical barriers -- slow mixing and imperfect quality control -- reveal watermarking to be far more robust than theoretical models suggest. The gap between idealized attacks and real-world feasibility underscores the need for stronger watermarking methods and more realistic attack models.
format Preprint
id arxiv_https___arxiv_org_abs_2505_06827
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sandcastles in the Storm: Revisiting the (Im)possibility of Strong Watermarking
Harel-Canada, Fabrice Y
Erol, Boran
Choi, Connor
Liu, Jason
Song, Gary Jiarui
Peng, Nanyun
Sahai, Amit
Cryptography and Security
Artificial Intelligence
Watermarking AI-generated text is critical for combating misuse. Yet recent theoretical work argues that any watermark can be erased via random walk attacks that perturb text while preserving quality. However, such attacks rely on two key assumptions: (1) rapid mixing (watermarks dissolve quickly under perturbations) and (2) reliable quality preservation (automated quality oracles perfectly guide edits). Through large-scale experiments and human-validated assessments, we find mixing is slow: 100% of perturbed texts retain traces of their origin after hundreds of edits, defying rapid mixing. Oracles falter, as state-of-the-art quality detectors misjudge edits (77% accuracy), compounding errors during attacks. Ultimately, attacks underperform: automated walks remove watermarks just 26% of the time -- dropping to 10% under human quality review. These findings challenge the inevitability of watermark removal. Instead, practical barriers -- slow mixing and imperfect quality control -- reveal watermarking to be far more robust than theoretical models suggest. The gap between idealized attacks and real-world feasibility underscores the need for stronger watermarking methods and more realistic attack models.
title Sandcastles in the Storm: Revisiting the (Im)possibility of Strong Watermarking
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2505.06827