Improving Detection of Watermarked Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bahri, Dara, Wieting, John
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918321751851008
author Bahri, Dara
Wieting, John
author_facet Bahri, Dara
Wieting, John
contents Watermarking has recently emerged as an effective strategy for detecting the generations of large language models (LLMs). The strength of a watermark typically depends strongly on the entropy afforded by the language model and the set of input prompts. However, entropy can be quite limited in practice, especially for models that are post-trained, for example via instruction tuning or reinforcement learning from human feedback (RLHF), which makes detection based on watermarking alone challenging. In this work, we investigate whether detection can be improved by combining watermark detectors with non-watermark ones. We explore a number of hybrid schemes that combine the two, observing performance gains over either class of detector under a wide range of experimental conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13131
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Detection of Watermarked Language Models
Bahri, Dara
Wieting, John
Computation and Language
Machine Learning
Watermarking has recently emerged as an effective strategy for detecting the generations of large language models (LLMs). The strength of a watermark typically depends strongly on the entropy afforded by the language model and the set of input prompts. However, entropy can be quite limited in practice, especially for models that are post-trained, for example via instruction tuning or reinforcement learning from human feedback (RLHF), which makes detection based on watermarking alone challenging. In this work, we investigate whether detection can be improved by combining watermark detectors with non-watermark ones. We explore a number of hybrid schemes that combine the two, observing performance gains over either class of detector under a wide range of experimental conditions.
title Improving Detection of Watermarked Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2508.13131