DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: An, Hyeseon, Park, Shinwoo, Woo, Suyeon, Han, Yo-Sub
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918350025654272
author An, Hyeseon
Park, Shinwoo
Woo, Suyeon
Han, Yo-Sub
author_facet An, Hyeseon
Park, Shinwoo
Woo, Suyeon
Han, Yo-Sub
contents The promise of LLM watermarking rests on a core assumption that a specific watermark proves authorship by a specific model. We demonstrate that this assumption is dangerously flawed. We introduce the threat of watermark spoofing, a sophisticated attack that allows a malicious model to generate text containing the authentic-looking watermark of a trusted, victim model. This enables the seamless misattribution of harmful content, such as disinformation, to reputable sources. The key to our attack is repurposing watermark radioactivity, the unintended inheritance of data patterns during fine-tuning, from a discoverable trait into an attack vector. By distilling knowledge from a watermarked teacher model, our framework allows an attacker to steal and replicate the watermarking signal of the victim model. This work reveals a critical security gap in text authorship verification and calls for a paradigm shift towards technologies capable of distinguishing authentic watermarks from expertly imitated ones. Our code is available at https://github.com/hsannn/ditto.git.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10987
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation
An, Hyeseon
Park, Shinwoo
Woo, Suyeon
Han, Yo-Sub
Cryptography and Security
Artificial Intelligence
The promise of LLM watermarking rests on a core assumption that a specific watermark proves authorship by a specific model. We demonstrate that this assumption is dangerously flawed. We introduce the threat of watermark spoofing, a sophisticated attack that allows a malicious model to generate text containing the authentic-looking watermark of a trusted, victim model. This enables the seamless misattribution of harmful content, such as disinformation, to reputable sources. The key to our attack is repurposing watermark radioactivity, the unintended inheritance of data patterns during fine-tuning, from a discoverable trait into an attack vector. By distilling knowledge from a watermarked teacher model, our framework allows an attacker to steal and replicate the watermarking signal of the victim model. This work reveals a critical security gap in text authorship verification and calls for a paradigm shift towards technologies capable of distinguishing authentic watermarks from expertly imitated ones. Our code is available at https://github.com/hsannn/ditto.git.
title DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2510.10987