WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tu, Shangqing, Sun, Yuliang, Bai, Yushi, Yu, Jifan, Hou, Lei, Li, Juanzi
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916306404507648
author Tu, Shangqing
Sun, Yuliang
Bai, Yushi
Yu, Jifan
Hou, Lei
Li, Juanzi
author_facet Tu, Shangqing
Sun, Yuliang
Bai, Yushi
Yu, Jifan
Hou, Lei
Li, Juanzi
contents To mitigate the potential misuse of large language models (LLMs), recent research has developed watermarking algorithms, which restrict the generation process to leave an invisible trace for watermark detection. Due to the two-stage nature of the task, most studies evaluate the generation and detection separately, thereby presenting a challenge in unbiased, thorough, and applicable evaluations. In this paper, we introduce WaterBench, the first comprehensive benchmark for LLM watermarks, in which we design three crucial factors: (1) For benchmarking procedure, to ensure an apples-to-apples comparison, we first adjust each watermarking method's hyper-parameter to reach the same watermarking strength, then jointly evaluate their generation and detection performance. (2) For task selection, we diversify the input and output length to form a five-category taxonomy, covering $9$ tasks. (3) For evaluation metric, we adopt the GPT4-Judge for automatically evaluating the decline of instruction-following abilities after watermarking. We evaluate $4$ open-source watermarks on $2$ LLMs under $2$ watermarking strengths and observe the common struggles for current methods on maintaining the generation quality. The code and data are available at https://github.com/THU-KEG/WaterBench.
format Preprint
id arxiv_https___arxiv_org_abs_2311_07138
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models
Tu, Shangqing
Sun, Yuliang
Bai, Yushi
Yu, Jifan
Hou, Lei
Li, Juanzi
Computation and Language
Artificial Intelligence
To mitigate the potential misuse of large language models (LLMs), recent research has developed watermarking algorithms, which restrict the generation process to leave an invisible trace for watermark detection. Due to the two-stage nature of the task, most studies evaluate the generation and detection separately, thereby presenting a challenge in unbiased, thorough, and applicable evaluations. In this paper, we introduce WaterBench, the first comprehensive benchmark for LLM watermarks, in which we design three crucial factors: (1) For benchmarking procedure, to ensure an apples-to-apples comparison, we first adjust each watermarking method's hyper-parameter to reach the same watermarking strength, then jointly evaluate their generation and detection performance. (2) For task selection, we diversify the input and output length to form a five-category taxonomy, covering $9$ tasks. (3) For evaluation metric, we adopt the GPT4-Judge for automatically evaluating the decline of instruction-following abilities after watermarking. We evaluate $4$ open-source watermarks on $2$ LLMs under $2$ watermarking strengths and observe the common struggles for current methods on maintaining the generation quality. The code and data are available at https://github.com/THU-KEG/WaterBench.
title WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2311.07138