Unified attacks to large language model watermarks: spoofing and scrubbing in unauthorized knowledge distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yi, Xin, Li, Yue, Zheng, Shunfan, Wang, Linlin, Wang, Xiaoling, He, Liang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909750474571776
author Yi, Xin
Li, Yue
Zheng, Shunfan
Wang, Linlin
Wang, Xiaoling
He, Liang
author_facet Yi, Xin
Li, Yue
Zheng, Shunfan
Wang, Linlin
Wang, Xiaoling
He, Liang
contents Watermarking has emerged as a critical technique for combating misinformation and protecting intellectual property in large language models (LLMs). A recent discovery, termed watermark radioactivity, reveals that watermarks embedded in teacher models can be inherited by student models through knowledge distillation. On the positive side, this inheritance allows for the detection of unauthorized knowledge distillation by identifying watermark traces in student models. However, the robustness of watermarks against scrubbing attacks and their unforgeability in the face of spoofing attacks under unauthorized knowledge distillation remain largely unexplored. Existing watermark attack methods either assume access to model internals or fail to simultaneously support both scrubbing and spoofing attacks. In this work, we propose Contrastive Decoding-Guided Knowledge Distillation (CDG-KD), a unified framework that enables bidirectional attacks under unauthorized knowledge distillation. Our approach employs contrastive decoding to extract corrupted or amplified watermark texts via comparing outputs from the student model and weakly watermarked references, followed by bidirectional distillation to train new student models capable of watermark removal and watermark forgery, respectively. Extensive experiments show that CDG-KD effectively performs attacks while preserving the general performance of the distilled model. Our findings underscore critical need for developing watermarking schemes that are robust and unforgeable.
format Preprint
id arxiv_https___arxiv_org_abs_2504_17480
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unified attacks to large language model watermarks: spoofing and scrubbing in unauthorized knowledge distillation
Yi, Xin
Li, Yue
Zheng, Shunfan
Wang, Linlin
Wang, Xiaoling
He, Liang
Computation and Language
Watermarking has emerged as a critical technique for combating misinformation and protecting intellectual property in large language models (LLMs). A recent discovery, termed watermark radioactivity, reveals that watermarks embedded in teacher models can be inherited by student models through knowledge distillation. On the positive side, this inheritance allows for the detection of unauthorized knowledge distillation by identifying watermark traces in student models. However, the robustness of watermarks against scrubbing attacks and their unforgeability in the face of spoofing attacks under unauthorized knowledge distillation remain largely unexplored. Existing watermark attack methods either assume access to model internals or fail to simultaneously support both scrubbing and spoofing attacks. In this work, we propose Contrastive Decoding-Guided Knowledge Distillation (CDG-KD), a unified framework that enables bidirectional attacks under unauthorized knowledge distillation. Our approach employs contrastive decoding to extract corrupted or amplified watermark texts via comparing outputs from the student model and weakly watermarked references, followed by bidirectional distillation to train new student models capable of watermark removal and watermark forgery, respectively. Extensive experiments show that CDG-KD effectively performs attacks while preserving the general performance of the distilled model. Our findings underscore critical need for developing watermarking schemes that are robust and unforgeable.
title Unified attacks to large language model watermarks: spoofing and scrubbing in unauthorized knowledge distillation
topic Computation and Language
url https://arxiv.org/abs/2504.17480