WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Molenda, Piotr, Liusie, Adian, Gales, Mark J. F.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911818528587776
author Molenda, Piotr
Liusie, Adian
Gales, Mark J. F.
author_facet Molenda, Piotr
Liusie, Adian
Gales, Mark J. F.
contents Watermarking generative-AI systems, such as LLMs, has gained considerable interest, driven by their enhanced capabilities across a wide range of tasks. Although current approaches have demonstrated that small, context-dependent shifts in the word distributions can be used to apply and detect watermarks, there has been little work in analyzing the impact that these perturbations have on the quality of generated texts. Balancing high detectability with minimal performance degradation is crucial in terms of selecting the appropriate watermarking setting; therefore this paper proposes a simple analysis framework where comparative assessment, a flexible NLG evaluation framework, is used to assess the quality degradation caused by a particular watermark setting. We demonstrate that our framework provides easy visualization of the quality-detection trade-off of watermark settings, enabling a simple solution to find an LLM watermark operating point that provides a well-balanced performance. This approach is applied to two different summarization systems and a translation system, enabling cross-model analysis for a task, and cross-task analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19548
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models
Molenda, Piotr
Liusie, Adian
Gales, Mark J. F.
Computation and Language
Watermarking generative-AI systems, such as LLMs, has gained considerable interest, driven by their enhanced capabilities across a wide range of tasks. Although current approaches have demonstrated that small, context-dependent shifts in the word distributions can be used to apply and detect watermarks, there has been little work in analyzing the impact that these perturbations have on the quality of generated texts. Balancing high detectability with minimal performance degradation is crucial in terms of selecting the appropriate watermarking setting; therefore this paper proposes a simple analysis framework where comparative assessment, a flexible NLG evaluation framework, is used to assess the quality degradation caused by a particular watermark setting. We demonstrate that our framework provides easy visualization of the quality-detection trade-off of watermark settings, enabling a simple solution to find an LLM watermark operating point that provides a well-balanced performance. This approach is applied to two different summarization systems and a translation system, enabling cross-model analysis for a task, and cross-task analysis.
title WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2403.19548