Towards Better Statistical Understanding of Watermarking LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cai, Zhongze, Liu, Shang, Wang, Hanzhao, Zhong, Huaiyang, Li, Xiaocheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914449309302784
author Cai, Zhongze
Liu, Shang
Wang, Hanzhao
Zhong, Huaiyang
Li, Xiaocheng
author_facet Cai, Zhongze
Liu, Shang
Wang, Hanzhao
Zhong, Huaiyang
Li, Xiaocheng
contents In this paper, we study the problem of watermarking large language models (LLMs). We consider the trade-off between model distortion and detection ability and formulate it as a constrained optimization problem based on the red-green list watermarking algorithm. We show that the optimal solution to the optimization problem enjoys a nice analytical property which provides a better understanding and inspires the algorithm design for the watermarking process. We develop an online dual gradient ascent watermarking algorithm in light of this optimization formulation and prove its asymptotic Pareto optimality between model distortion and detection ability. Such a result guarantees an averaged increased green list probability and henceforth detection ability explicitly (in contrast to previous results). Moreover, we provide a systematic discussion on the choice of the model distortion metrics for the watermarking problem. We justify our choice of KL divergence and present issues with the existing criteria of ``distortion-free'' and perplexity. Finally, we empirically evaluate our algorithms on extensive datasets against benchmark algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2403_13027
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Better Statistical Understanding of Watermarking LLMs
Cai, Zhongze
Liu, Shang
Wang, Hanzhao
Zhong, Huaiyang
Li, Xiaocheng
Machine Learning
Cryptography and Security
Information Theory
In this paper, we study the problem of watermarking large language models (LLMs). We consider the trade-off between model distortion and detection ability and formulate it as a constrained optimization problem based on the red-green list watermarking algorithm. We show that the optimal solution to the optimization problem enjoys a nice analytical property which provides a better understanding and inspires the algorithm design for the watermarking process. We develop an online dual gradient ascent watermarking algorithm in light of this optimization formulation and prove its asymptotic Pareto optimality between model distortion and detection ability. Such a result guarantees an averaged increased green list probability and henceforth detection ability explicitly (in contrast to previous results). Moreover, we provide a systematic discussion on the choice of the model distortion metrics for the watermarking problem. We justify our choice of KL divergence and present issues with the existing criteria of ``distortion-free'' and perplexity. Finally, we empirically evaluate our algorithms on extensive datasets against benchmark algorithms.
title Towards Better Statistical Understanding of Watermarking LLMs
topic Machine Learning
Cryptography and Security
Information Theory
url https://arxiv.org/abs/2403.13027