Perturb Your Data: Paraphrase-Guided Training Data Watermarking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shetty, Pranav, Haque, Mirazul, Babkin, Petr, Ma, Zhiqiang, Liu, Xiaomo, Veloso, Manuela
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908908065390592
author Shetty, Pranav
Haque, Mirazul
Babkin, Petr
Ma, Zhiqiang
Liu, Xiaomo
Veloso, Manuela
author_facet Shetty, Pranav
Haque, Mirazul
Babkin, Petr
Ma, Zhiqiang
Liu, Xiaomo
Veloso, Manuela
contents Training data detection is critical for enforcing copyright and data licensing, as Large Language Models (LLM) are trained on massive text corpora scraped from the internet. We present SPECTRA, a watermarking approach that makes training data reliably detectable even when it comprises less than 0.001% of the training corpus. SPECTRA works by paraphrasing text using an LLM and assigning a score based on how likely each paraphrase is, according to a separate scoring model. A paraphrase is chosen so that its score closely matches that of the original text, to avoid introducing any distribution shifts. To test whether a suspect model has been trained on the watermarked data, we compare its token probabilities against those of the scoring model. We demonstrate that SPECTRA achieves a consistent p-value gap of over nine orders of magnitude when detecting data used for training versus data not used for training, which is greater than all baselines tested. SPECTRA equips data owners with a scalable, deploy-before-release watermark that survives even large-scale LLM training.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17075
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Perturb Your Data: Paraphrase-Guided Training Data Watermarking
Shetty, Pranav
Haque, Mirazul
Babkin, Petr
Ma, Zhiqiang
Liu, Xiaomo
Veloso, Manuela
Computation and Language
Machine Learning
Training data detection is critical for enforcing copyright and data licensing, as Large Language Models (LLM) are trained on massive text corpora scraped from the internet. We present SPECTRA, a watermarking approach that makes training data reliably detectable even when it comprises less than 0.001% of the training corpus. SPECTRA works by paraphrasing text using an LLM and assigning a score based on how likely each paraphrase is, according to a separate scoring model. A paraphrase is chosen so that its score closely matches that of the original text, to avoid introducing any distribution shifts. To test whether a suspect model has been trained on the watermarked data, we compare its token probabilities against those of the scoring model. We demonstrate that SPECTRA achieves a consistent p-value gap of over nine orders of magnitude when detecting data used for training versus data not used for training, which is greater than all baselines tested. SPECTRA equips data owners with a scalable, deploy-before-release watermark that survives even large-scale LLM training.
title Perturb Your Data: Paraphrase-Guided Training Data Watermarking
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2512.17075