Advancing Beyond Identification: Multi-bit Watermark for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yoo, KiYoon, Ahn, Wonhyuk, Kwak, Nojun
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911803983790080
author Yoo, KiYoon
Ahn, Wonhyuk
Kwak, Nojun
author_facet Yoo, KiYoon
Ahn, Wonhyuk
Kwak, Nojun
contents We show the viability of tackling misuses of large language models beyond the identification of machine-generated text. While existing zero-bit watermark methods focus on detection only, some malicious misuses demand tracing the adversary user for counteracting them. To address this, we propose Multi-bit Watermark via Position Allocation, embedding traceable multi-bit information during language model generation. Through allocating tokens onto different parts of the messages, we embed longer messages in high corruption settings without added latency. By independently embedding sub-units of messages, the proposed method outperforms the existing works in terms of robustness and latency. Leveraging the benefits of zero-bit watermarking, our method enables robust extraction of the watermark without any model access, embedding and extraction of long messages ($\geq$ 32-bit) without finetuning, and maintaining text quality, while allowing zero-bit detection all at the same time. Code is released here: https://github.com/bangawayoo/mb-lm-watermarking
format Preprint
id arxiv_https___arxiv_org_abs_2308_00221
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Advancing Beyond Identification: Multi-bit Watermark for Large Language Models
Yoo, KiYoon
Ahn, Wonhyuk
Kwak, Nojun
Computation and Language
Artificial Intelligence
Cryptography and Security
We show the viability of tackling misuses of large language models beyond the identification of machine-generated text. While existing zero-bit watermark methods focus on detection only, some malicious misuses demand tracing the adversary user for counteracting them. To address this, we propose Multi-bit Watermark via Position Allocation, embedding traceable multi-bit information during language model generation. Through allocating tokens onto different parts of the messages, we embed longer messages in high corruption settings without added latency. By independently embedding sub-units of messages, the proposed method outperforms the existing works in terms of robustness and latency. Leveraging the benefits of zero-bit watermarking, our method enables robust extraction of the watermark without any model access, embedding and extraction of long messages ($\geq$ 32-bit) without finetuning, and maintaining text quality, while allowing zero-bit detection all at the same time. Code is released here: https://github.com/bangawayoo/mb-lm-watermarking
title Advancing Beyond Identification: Multi-bit Watermark for Large Language Models
topic Computation and Language
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2308.00221