Towards Watermarking of Open-Source LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gloaguen, Thibaud, Jovanović, Nikola, Staab, Robin, Vechev, Martin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910828583714816
author Gloaguen, Thibaud
Jovanović, Nikola
Staab, Robin
Vechev, Martin
author_facet Gloaguen, Thibaud
Jovanović, Nikola
Staab, Robin
Vechev, Martin
contents While watermarks for closed LLMs have matured and have been included in large-scale deployments, these methods are not applicable to open-source models, which allow users full control over the decoding process. This setting is understudied yet critical, given the rising performance of open-source models. In this work, we lay the foundation for systematic study of open-source LLM watermarking. For the first time, we explicitly formulate key requirements, including durability against common model modifications such as model merging, quantization, or finetuning, and propose a concrete evaluation setup. Given the prevalence of these modifications, durability is crucial for an open-source watermark to be effective. We survey and evaluate existing methods, showing that they are not durable. We also discuss potential ways to improve their durability and highlight remaining challenges. We hope our work enables future progress on this important problem.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10525
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Watermarking of Open-Source LLMs
Gloaguen, Thibaud
Jovanović, Nikola
Staab, Robin
Vechev, Martin
Cryptography and Security
Machine Learning
While watermarks for closed LLMs have matured and have been included in large-scale deployments, these methods are not applicable to open-source models, which allow users full control over the decoding process. This setting is understudied yet critical, given the rising performance of open-source models. In this work, we lay the foundation for systematic study of open-source LLM watermarking. For the first time, we explicitly formulate key requirements, including durability against common model modifications such as model merging, quantization, or finetuning, and propose a concrete evaluation setup. Given the prevalence of these modifications, durability is crucial for an open-source watermark to be effective. We survey and evaluate existing methods, showing that they are not durable. We also discuss potential ways to improve their durability and highlight remaining challenges. We hope our work enables future progress on this important problem.
title Towards Watermarking of Open-Source LLMs
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2502.10525