Robust Multi-bit Text Watermark with LLM-based Paraphrasers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Xiaojun, Jia, Jinghan, Yao, Yuanshun, Liu, Yang, Li, Hang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909650278940672
author Xu, Xiaojun
Jia, Jinghan
Yao, Yuanshun
Liu, Yang
Li, Hang
author_facet Xu, Xiaojun
Jia, Jinghan
Yao, Yuanshun
Liu, Yang
Li, Hang
contents We propose an imperceptible multi-bit text watermark embedded by paraphrasing with LLMs. We fine-tune a pair of LLM paraphrasers that are designed to behave differently so that their paraphrasing difference reflected in the text semantics can be identified by a trained decoder. To embed our multi-bit watermark, we use two paraphrasers alternatively to encode the pre-defined binary code at the sentence level. Then we use a text classifier as the decoder to decode each bit of the watermark. Through extensive experiments, we show that our watermarks can achieve over 99.99\% detection AUC with small (1.1B) text paraphrasers while keeping the semantic information of the original sentence. More importantly, our pipeline is robust under word substitution and sentence paraphrasing perturbations and generalizes well to out-of-distributional data. We also show the stealthiness of our watermark with LLM-based evaluation. We open-source the code: https://github.com/xiaojunxu/multi-bit-text-watermark.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03123
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Robust Multi-bit Text Watermark with LLM-based Paraphrasers
Xu, Xiaojun
Jia, Jinghan
Yao, Yuanshun
Liu, Yang
Li, Hang
Artificial Intelligence
We propose an imperceptible multi-bit text watermark embedded by paraphrasing with LLMs. We fine-tune a pair of LLM paraphrasers that are designed to behave differently so that their paraphrasing difference reflected in the text semantics can be identified by a trained decoder. To embed our multi-bit watermark, we use two paraphrasers alternatively to encode the pre-defined binary code at the sentence level. Then we use a text classifier as the decoder to decode each bit of the watermark. Through extensive experiments, we show that our watermarks can achieve over 99.99\% detection AUC with small (1.1B) text paraphrasers while keeping the semantic information of the original sentence. More importantly, our pipeline is robust under word substitution and sentence paraphrasing perturbations and generalizes well to out-of-distributional data. We also show the stealthiness of our watermark with LLM-based evaluation. We open-source the code: https://github.com/xiaojunxu/multi-bit-text-watermark.
title Robust Multi-bit Text Watermark with LLM-based Paraphrasers
topic Artificial Intelligence
url https://arxiv.org/abs/2412.03123