Revisiting the Robustness of Watermarking to Paraphrasing Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rastogi, Saksham, Pruthi, Danish
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917831557251072
author Rastogi, Saksham
Pruthi, Danish
author_facet Rastogi, Saksham
Pruthi, Danish
contents Amidst rising concerns about the internet being proliferated with content generated from language models (LMs), watermarking is seen as a principled way to certify whether text was generated from a model. Many recent watermarking techniques slightly modify the output probabilities of LMs to embed a signal in the generated output that can later be detected. Since early proposals for text watermarking, questions about their robustness to paraphrasing have been prominently discussed. Lately, some techniques are deliberately designed and claimed to be robust to paraphrasing. However, such watermarking schemes do not adequately account for the ease with which they can be reverse-engineered. We show that with access to only a limited number of generations from a black-box watermarked model, we can drastically increase the effectiveness of paraphrasing attacks to evade watermark detection, thereby rendering the watermark ineffective.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05277
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Revisiting the Robustness of Watermarking to Paraphrasing Attacks
Rastogi, Saksham
Pruthi, Danish
Cryptography and Security
Computation and Language
Machine Learning
Amidst rising concerns about the internet being proliferated with content generated from language models (LMs), watermarking is seen as a principled way to certify whether text was generated from a model. Many recent watermarking techniques slightly modify the output probabilities of LMs to embed a signal in the generated output that can later be detected. Since early proposals for text watermarking, questions about their robustness to paraphrasing have been prominently discussed. Lately, some techniques are deliberately designed and claimed to be robust to paraphrasing. However, such watermarking schemes do not adequately account for the ease with which they can be reverse-engineered. We show that with access to only a limited number of generations from a black-box watermarked model, we can drastically increase the effectiveness of paraphrasing attacks to evade watermark detection, thereby rendering the watermark ineffective.
title Revisiting the Robustness of Watermarking to Paraphrasing Attacks
topic Cryptography and Security
Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.05277