WAPITI: A Watermark for Finetuned Open-Source LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Lingjie, Qiu, Ruizhong, Yuan, Siyu, Liu, Zhining, Wei, Tianxin, Yoo, Hyunsik, Zeng, Zhichen, Yang, Deqing, Tong, Hanghang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917859361292288
author Chen, Lingjie
Qiu, Ruizhong
Yuan, Siyu
Liu, Zhining
Wei, Tianxin
Yoo, Hyunsik
Zeng, Zhichen
Yang, Deqing
Tong, Hanghang
author_facet Chen, Lingjie
Qiu, Ruizhong
Yuan, Siyu
Liu, Zhining
Wei, Tianxin
Yoo, Hyunsik
Zeng, Zhichen
Yang, Deqing
Tong, Hanghang
contents Watermarking of large language models (LLMs) generation embeds an imperceptible statistical pattern within texts, making it algorithmically detectable. Watermarking is a promising method for addressing potential harm and biases from LLMs, as it enables traceability, accountability, and detection of manipulated content, helping to mitigate unintended consequences. However, for open-source models, watermarking faces two major challenges: (i) incompatibility with fine-tuned models, and (ii) vulnerability to fine-tuning attacks. In this work, we propose WAPITI, a new method that transfers watermarking from base models to fine-tuned models through parameter integration. To the best of our knowledge, we propose the first watermark for fine-tuned open-source LLMs that preserves their fine-tuned capabilities. Furthermore, our approach offers an effective defense against fine-tuning attacks. We test our method on various model architectures and watermarking strategies. Results demonstrate that our method can successfully inject watermarks and is highly compatible with fine-tuned models. Additionally, we offer an in-depth analysis of how parameter editing influences the watermark strength and overall capabilities of the resulting models.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06467
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WAPITI: A Watermark for Finetuned Open-Source LLMs
Chen, Lingjie
Qiu, Ruizhong
Yuan, Siyu
Liu, Zhining
Wei, Tianxin
Yoo, Hyunsik
Zeng, Zhichen
Yang, Deqing
Tong, Hanghang
Cryptography and Security
Watermarking of large language models (LLMs) generation embeds an imperceptible statistical pattern within texts, making it algorithmically detectable. Watermarking is a promising method for addressing potential harm and biases from LLMs, as it enables traceability, accountability, and detection of manipulated content, helping to mitigate unintended consequences. However, for open-source models, watermarking faces two major challenges: (i) incompatibility with fine-tuned models, and (ii) vulnerability to fine-tuning attacks. In this work, we propose WAPITI, a new method that transfers watermarking from base models to fine-tuned models through parameter integration. To the best of our knowledge, we propose the first watermark for fine-tuned open-source LLMs that preserves their fine-tuned capabilities. Furthermore, our approach offers an effective defense against fine-tuning attacks. We test our method on various model architectures and watermarking strategies. Results demonstrate that our method can successfully inject watermarks and is highly compatible with fine-tuned models. Additionally, we offer an in-depth analysis of how parameter editing influences the watermark strength and overall capabilities of the resulting models.
title WAPITI: A Watermark for Finetuned Open-Source LLMs
topic Cryptography and Security
url https://arxiv.org/abs/2410.06467