Breaking Distortion-free Watermarks in Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Reynolds, Shayleen, He, Hengzhi, Ngo, Dung Daniel T., Obitayo, Saheed, Dalmasso, Niccolò, Cheng, Guang, Potluru, Vamsi K., Veloso, Manuela
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915338613948416
author Reynolds, Shayleen
He, Hengzhi
Ngo, Dung Daniel T.
Obitayo, Saheed
Dalmasso, Niccolò
Cheng, Guang
Potluru, Vamsi K.
Veloso, Manuela
author_facet Reynolds, Shayleen
He, Hengzhi
Ngo, Dung Daniel T.
Obitayo, Saheed
Dalmasso, Niccolò
Cheng, Guang
Potluru, Vamsi K.
Veloso, Manuela
contents In recent years, LLM watermarking has emerged as an attractive safeguard against AI-generated content, with promising applications in many real-world domains. However, there are growing concerns that the current LLM watermarking schemes are vulnerable to expert adversaries wishing to reverse-engineer the watermarking mechanisms. Prior work in breaking or stealing LLM watermarks mainly focuses on the distribution-modifying algorithm of Kirchenbauer et al. (2023), which perturbs the logit vector before sampling. In this work, we focus on reverse-engineering the other prominent LLM watermarking scheme, distortion-free watermarking (Kuditipudi et al. 2024), which preserves the underlying token distribution by using a hidden watermarking key sequence. We demonstrate that, even under a more sophisticated watermarking scheme, it is possible to compromise the LLM and carry out a spoofing attack, i.e. generate a large number of (potentially harmful) texts that can be attributed to the original watermarked LLM. Specifically, we propose using adaptive prompting and a sorting-based algorithm to accurately recover the underlying secret key for watermarking the LLM. Our empirical findings on LLAMA-3.1-8B-Instruct, Mistral-7B-Instruct, Gemma-7b, and OPT-125M challenge the current theoretical claims on the robustness and usability of the distortion-free watermarking techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2502_18608
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Breaking Distortion-free Watermarks in Large Language Models
Reynolds, Shayleen
He, Hengzhi
Ngo, Dung Daniel T.
Obitayo, Saheed
Dalmasso, Niccolò
Cheng, Guang
Potluru, Vamsi K.
Veloso, Manuela
Cryptography and Security
Machine Learning
In recent years, LLM watermarking has emerged as an attractive safeguard against AI-generated content, with promising applications in many real-world domains. However, there are growing concerns that the current LLM watermarking schemes are vulnerable to expert adversaries wishing to reverse-engineer the watermarking mechanisms. Prior work in breaking or stealing LLM watermarks mainly focuses on the distribution-modifying algorithm of Kirchenbauer et al. (2023), which perturbs the logit vector before sampling. In this work, we focus on reverse-engineering the other prominent LLM watermarking scheme, distortion-free watermarking (Kuditipudi et al. 2024), which preserves the underlying token distribution by using a hidden watermarking key sequence. We demonstrate that, even under a more sophisticated watermarking scheme, it is possible to compromise the LLM and carry out a spoofing attack, i.e. generate a large number of (potentially harmful) texts that can be attributed to the original watermarked LLM. Specifically, we propose using adaptive prompting and a sorting-based algorithm to accurately recover the underlying secret key for watermarking the LLM. Our empirical findings on LLAMA-3.1-8B-Instruct, Mistral-7B-Instruct, Gemma-7b, and OPT-125M challenge the current theoretical claims on the robustness and usability of the distortion-free watermarking techniques.
title Breaking Distortion-free Watermarks in Large Language Models
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2502.18608