SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910416405266432 |
|---|---|
| author | Hou, Abe Bohan Zhang, Jingyu He, Tianxing Wang, Yichen Chuang, Yung-Sung Wang, Hongwei Shen, Lingfeng Van Durme, Benjamin Khashabi, Daniel Tsvetkov, Yulia |
| author_facet | Hou, Abe Bohan Zhang, Jingyu He, Tianxing Wang, Yichen Chuang, Yung-Sung Wang, Hongwei Shen, Lingfeng Van Durme, Benjamin Khashabi, Daniel Tsvetkov, Yulia |
| contents | Existing watermarking algorithms are vulnerable to paraphrase attacks because of their token-level design. To address this issue, we propose SemStamp, a robust sentence-level semantic watermarking algorithm based on locality-sensitive hashing (LSH), which partitions the semantic space of sentences. The algorithm encodes and LSH-hashes a candidate sentence generated by an LLM, and conducts sentence-level rejection sampling until the sampled sentence falls in watermarked partitions in the semantic embedding space. A margin-based constraint is used to enhance its robustness. To show the advantages of our algorithm, we propose a "bigram" paraphrase attack using the paraphrase that has the fewest bigram overlaps with the original sentence. This attack is shown to be effective against the existing token-level watermarking method. Experimental results show that our novel semantic watermark algorithm is not only more robust than the previous state-of-the-art method on both common and bigram paraphrase attacks, but also is better at preserving the quality of generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2310_03991 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation Hou, Abe Bohan Zhang, Jingyu He, Tianxing Wang, Yichen Chuang, Yung-Sung Wang, Hongwei Shen, Lingfeng Van Durme, Benjamin Khashabi, Daniel Tsvetkov, Yulia Computation and Language Existing watermarking algorithms are vulnerable to paraphrase attacks because of their token-level design. To address this issue, we propose SemStamp, a robust sentence-level semantic watermarking algorithm based on locality-sensitive hashing (LSH), which partitions the semantic space of sentences. The algorithm encodes and LSH-hashes a candidate sentence generated by an LLM, and conducts sentence-level rejection sampling until the sampled sentence falls in watermarked partitions in the semantic embedding space. A margin-based constraint is used to enhance its robustness. To show the advantages of our algorithm, we propose a "bigram" paraphrase attack using the paraphrase that has the fewest bigram overlaps with the original sentence. This attack is shown to be effective against the existing token-level watermarking method. Experimental results show that our novel semantic watermark algorithm is not only more robust than the previous state-of-the-art method on both common and bigram paraphrase attacks, but also is better at preserving the quality of generation. |
| title | SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2310.03991 |