Watermarking Needs Input Repetition Masking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khachaturov, David, Mullins, Robert, Shumailov, Ilia, Dathathri, Sumanth
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912331547541504
author Khachaturov, David
Mullins, Robert
Shumailov, Ilia
Dathathri, Sumanth
author_facet Khachaturov, David
Mullins, Robert
Shumailov, Ilia
Dathathri, Sumanth
contents Recent advancements in Large Language Models (LLMs) raised concerns over potential misuse, such as for spreading misinformation. In response two counter measures emerged: machine learning-based detectors that predict if text is synthetic, and LLM watermarking, which subtly marks generated text for identification and attribution. Meanwhile, humans are known to adjust language to their conversational partners both syntactically and lexically. By implication, it is possible that humans or unwatermarked LLMs could unintentionally mimic properties of LLM generated text, making counter measures unreliable. In this work we investigate the extent to which such conversational adaptation happens. We call the concept $\textit{mimicry}$ and demonstrate that both humans and LLMs end up mimicking, including the watermarking signal even in seemingly improbable settings. This challenges current academic assumptions and suggests that for long-term watermarking to be reliable, the likelihood of false positives needs to be significantly lower, while longer word sequences should be used for seeding watermarking mechanisms.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12229
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Watermarking Needs Input Repetition Masking
Khachaturov, David
Mullins, Robert
Shumailov, Ilia
Dathathri, Sumanth
Machine Learning
Computation and Language
Cryptography and Security
Recent advancements in Large Language Models (LLMs) raised concerns over potential misuse, such as for spreading misinformation. In response two counter measures emerged: machine learning-based detectors that predict if text is synthetic, and LLM watermarking, which subtly marks generated text for identification and attribution. Meanwhile, humans are known to adjust language to their conversational partners both syntactically and lexically. By implication, it is possible that humans or unwatermarked LLMs could unintentionally mimic properties of LLM generated text, making counter measures unreliable. In this work we investigate the extent to which such conversational adaptation happens. We call the concept $\textit{mimicry}$ and demonstrate that both humans and LLMs end up mimicking, including the watermarking signal even in seemingly improbable settings. This challenges current academic assumptions and suggests that for long-term watermarking to be reliable, the likelihood of false positives needs to be significantly lower, while longer word sequences should be used for seeding watermarking mechanisms.
title Watermarking Needs Input Repetition Masking
topic Machine Learning
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2504.12229