Learning to Watermark LLM-generated Text via Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Xiaojun, Yao, Yuanshun, Liu, Yang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911799080648704
author Xu, Xiaojun
Yao, Yuanshun
Liu, Yang
author_facet Xu, Xiaojun
Yao, Yuanshun
Liu, Yang
contents We study how to watermark LLM outputs, i.e. embedding algorithmically detectable signals into LLM-generated text to track misuse. Unlike the current mainstream methods that work with a fixed LLM, we expand the watermark design space by including the LLM tuning stage in the watermark pipeline. While prior works focus on token-level watermark that embeds signals into the output, we design a model-level watermark that embeds signals into the LLM weights, and such signals can be detected by a paired detector. We propose a co-training framework based on reinforcement learning that iteratively (1) trains a detector to detect the generated watermarked text and (2) tunes the LLM to generate text easily detectable by the detector while keeping its normal utility. We empirically show that our watermarks are more accurate, robust, and adaptable (to new attacks). It also allows watermarked model open-sourcing. In addition, if used together with alignment, the extra overhead introduced is low - only training an extra reward model (i.e. our detector). We hope our work can bring more effort into studying a broader watermark design that is not limited to working with a fixed LLM. We open-source the code: https://github.com/xiaojunxu/learning-to-watermark-llm .
format Preprint
id arxiv_https___arxiv_org_abs_2403_10553
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning to Watermark LLM-generated Text via Reinforcement Learning
Xu, Xiaojun
Yao, Yuanshun
Liu, Yang
Machine Learning
Artificial Intelligence
Cryptography and Security
We study how to watermark LLM outputs, i.e. embedding algorithmically detectable signals into LLM-generated text to track misuse. Unlike the current mainstream methods that work with a fixed LLM, we expand the watermark design space by including the LLM tuning stage in the watermark pipeline. While prior works focus on token-level watermark that embeds signals into the output, we design a model-level watermark that embeds signals into the LLM weights, and such signals can be detected by a paired detector. We propose a co-training framework based on reinforcement learning that iteratively (1) trains a detector to detect the generated watermarked text and (2) tunes the LLM to generate text easily detectable by the detector while keeping its normal utility. We empirically show that our watermarks are more accurate, robust, and adaptable (to new attacks). It also allows watermarked model open-sourcing. In addition, if used together with alignment, the extra overhead introduced is low - only training an extra reward model (i.e. our detector). We hope our work can bring more effort into studying a broader watermark design that is not limited to working with a fixed LLM. We open-source the code: https://github.com/xiaojunxu/learning-to-watermark-llm .
title Learning to Watermark LLM-generated Text via Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2403.10553