Large Language Model Watermark Stealing With Mixed Integer Programming

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhaoxi, Zhang, Xiaomei, Zhang, Yanjun, Zhang, Leo Yu, Chen, Chao, Hu, Shengshan, Gill, Asif, Pan, Shirui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911894162374656
author Zhang, Zhaoxi
Zhang, Xiaomei
Zhang, Yanjun
Zhang, Leo Yu
Chen, Chao
Hu, Shengshan
Gill, Asif
Pan, Shirui
author_facet Zhang, Zhaoxi
Zhang, Xiaomei
Zhang, Yanjun
Zhang, Leo Yu
Chen, Chao
Hu, Shengshan
Gill, Asif
Pan, Shirui
contents The Large Language Model (LLM) watermark is a newly emerging technique that shows promise in addressing concerns surrounding LLM copyright, monitoring AI-generated text, and preventing its misuse. The LLM watermark scheme commonly includes generating secret keys to partition the vocabulary into green and red lists, applying a perturbation to the logits of tokens in the green list to increase their sampling likelihood, thus facilitating watermark detection to identify AI-generated text if the proportion of green tokens exceeds a threshold. However, recent research indicates that watermarking methods using numerous keys are susceptible to removal attacks, such as token editing, synonym substitution, and paraphrasing, with robustness declining as the number of keys increases. Therefore, the state-of-the-art watermark schemes that employ fewer or single keys have been demonstrated to be more robust against text editing and paraphrasing. In this paper, we propose a novel green list stealing attack against the state-of-the-art LLM watermark scheme and systematically examine its vulnerability to this attack. We formalize the attack as a mixed integer programming problem with constraints. We evaluate our attack under a comprehensive threat model, including an extreme scenario where the attacker has no prior knowledge, lacks access to the watermark detector API, and possesses no information about the LLM's parameter settings or watermark injection/detection scheme. Extensive experiments on LLMs, such as OPT and LLaMA, demonstrate that our attack can successfully steal the green list and remove the watermark across all settings.
format Preprint
id arxiv_https___arxiv_org_abs_2405_19677
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large Language Model Watermark Stealing With Mixed Integer Programming
Zhang, Zhaoxi
Zhang, Xiaomei
Zhang, Yanjun
Zhang, Leo Yu
Chen, Chao
Hu, Shengshan
Gill, Asif
Pan, Shirui
Cryptography and Security
Artificial Intelligence
The Large Language Model (LLM) watermark is a newly emerging technique that shows promise in addressing concerns surrounding LLM copyright, monitoring AI-generated text, and preventing its misuse. The LLM watermark scheme commonly includes generating secret keys to partition the vocabulary into green and red lists, applying a perturbation to the logits of tokens in the green list to increase their sampling likelihood, thus facilitating watermark detection to identify AI-generated text if the proportion of green tokens exceeds a threshold. However, recent research indicates that watermarking methods using numerous keys are susceptible to removal attacks, such as token editing, synonym substitution, and paraphrasing, with robustness declining as the number of keys increases. Therefore, the state-of-the-art watermark schemes that employ fewer or single keys have been demonstrated to be more robust against text editing and paraphrasing. In this paper, we propose a novel green list stealing attack against the state-of-the-art LLM watermark scheme and systematically examine its vulnerability to this attack. We formalize the attack as a mixed integer programming problem with constraints. We evaluate our attack under a comprehensive threat model, including an extreme scenario where the attacker has no prior knowledge, lacks access to the watermark detector API, and possesses no information about the LLM's parameter settings or watermark injection/detection scheme. Extensive experiments on LLMs, such as OPT and LLaMA, demonstrate that our attack can successfully steal the green list and remove the watermark across all settings.
title Large Language Model Watermark Stealing With Mixed Integer Programming
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2405.19677