Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Elhassan, Fay, Ajroldi, Niccolò, Orvieto, Antonio, Geiping, Jonas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909572122279936
author Elhassan, Fay
Ajroldi, Niccolò
Orvieto, Antonio
Geiping, Jonas
author_facet Elhassan, Fay
Ajroldi, Niccolò
Orvieto, Antonio
Geiping, Jonas
contents The indistinguishability of AI-generated content from human text raises challenges in transparency and accountability. While several methods exist to watermark models behind APIs, embedding watermark strategies directly into model weights that are later reflected in the outputs of the model is challenging. In this study we propose a strategy to finetune a pair of low-rank adapters of a model, one serving as the text-generating model, and the other as the detector, so that a subtle watermark is embedded into the text generated by the first model and simultaneously optimized for detectability by the second. In this way, the watermarking strategy is fully learned end-to-end. This process imposes an optimization challenge, as balancing watermark robustness, naturalness, and task performance requires trade-offs. We discuss strategies on how to optimize this min-max objective and present results showing the effect of this modification to instruction finetuning.
format Preprint
id arxiv_https___arxiv_org_abs_2504_06446
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
Elhassan, Fay
Ajroldi, Niccolò
Orvieto, Antonio
Geiping, Jonas
Machine Learning
Artificial Intelligence
The indistinguishability of AI-generated content from human text raises challenges in transparency and accountability. While several methods exist to watermark models behind APIs, embedding watermark strategies directly into model weights that are later reflected in the outputs of the model is challenging. In this study we propose a strategy to finetune a pair of low-rank adapters of a model, one serving as the text-generating model, and the other as the detector, so that a subtle watermark is embedded into the text generated by the first model and simultaneously optimized for detectability by the second. In this way, the watermarking strategy is fully learned end-to-end. This process imposes an optimization challenge, as balancing watermark robustness, naturalness, and task performance requires trade-offs. We discuss strategies on how to optimize this min-max objective and present results showing the effect of this modification to instruction finetuning.
title Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.06446