Multi-Bit Distortion-Free Watermarking for Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Boroujeny, Massieh Kordi, Jiang, Ya, Zeng, Kai, Mark, Brian
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929256767946752
author Boroujeny, Massieh Kordi
Jiang, Ya
Zeng, Kai
Mark, Brian
author_facet Boroujeny, Massieh Kordi
Jiang, Ya
Zeng, Kai
Mark, Brian
contents Methods for watermarking large language models have been proposed that distinguish AI-generated text from human-generated text by slightly altering the model output distribution, but they also distort the quality of the text, exposing the watermark to adversarial detection. More recently, distortion-free watermarking methods were proposed that require a secret key to detect the watermark. The prior methods generally embed zero-bit watermarks that do not provide additional information beyond tagging a text as being AI-generated. We extend an existing zero-bit distortion-free watermarking method by embedding multiple bits of meta-information as part of the watermark. We also develop a computationally efficient decoder that extracts the embedded information from the watermark with low bit error rate.
format Preprint
id arxiv_https___arxiv_org_abs_2402_16578
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Bit Distortion-Free Watermarking for Large Language Models
Boroujeny, Massieh Kordi
Jiang, Ya
Zeng, Kai
Mark, Brian
Computation and Language
Machine Learning
Methods for watermarking large language models have been proposed that distinguish AI-generated text from human-generated text by slightly altering the model output distribution, but they also distort the quality of the text, exposing the watermark to adversarial detection. More recently, distortion-free watermarking methods were proposed that require a secret key to detect the watermark. The prior methods generally embed zero-bit watermarks that do not provide additional information beyond tagging a text as being AI-generated. We extend an existing zero-bit distortion-free watermarking method by embedding multiple bits of meta-information as part of the watermark. We also develop a computationally efficient decoder that extracts the embedded information from the watermark with low bit error rate.
title Multi-Bit Distortion-Free Watermarking for Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2402.16578