SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918521466781696 |
|---|---|
| author | Huo, Jiahao Qu, Wenjie Yan, Yibo Zheng, Kening Zhang, Jiaheng Hu, Xuming Yu, Philip S. Zhou, Mingxun |
| author_facet | Huo, Jiahao Qu, Wenjie Yan, Yibo Zheng, Kening Zhang, Jiaheng Hu, Xuming Yu, Philip S. Zhou, Mingxun |
| contents | Semantic-level watermarking (SWM) improves robustness against text modifications by treating sentences as the basic unit. However, robustness to paragraph-level paraphrasing remains difficult because such attacks globally disrupt watermark signals by changing sentence order. In this work, we propose SAMark, a self-anchored watermarking framework that removes the dependency on sentence order by establishing a step-independent green region in semantic space. To improve detectability, we introduce a multi-channel hyperbolic scoring mechanism that amplifies watermark signals while suppressing noise from weakly aligned candidates. We further propose a diversity-aware filtering strategy that combines hard filtering with soft regularization, extending beyond simple n-gram repetition filters to address semantic redundancy. Experimental results show that SAMark achieves up to 90.2% TP@FP1% under typical paragraph-level paraphrasing attacks, outperforming the strongest prior baseline by more than 30% on average, while maintaining generation quality competitive with unwatermarked text and breaking the robustness-quality trade-off that limits prior methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_25796 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness Huo, Jiahao Qu, Wenjie Yan, Yibo Zheng, Kening Zhang, Jiaheng Hu, Xuming Yu, Philip S. Zhou, Mingxun Cryptography and Security Artificial Intelligence Computation and Language Semantic-level watermarking (SWM) improves robustness against text modifications by treating sentences as the basic unit. However, robustness to paragraph-level paraphrasing remains difficult because such attacks globally disrupt watermark signals by changing sentence order. In this work, we propose SAMark, a self-anchored watermarking framework that removes the dependency on sentence order by establishing a step-independent green region in semantic space. To improve detectability, we introduce a multi-channel hyperbolic scoring mechanism that amplifies watermark signals while suppressing noise from weakly aligned candidates. We further propose a diversity-aware filtering strategy that combines hard filtering with soft regularization, extending beyond simple n-gram repetition filters to address semantic redundancy. Experimental results show that SAMark achieves up to 90.2% TP@FP1% under typical paragraph-level paraphrasing attacks, outperforming the strongest prior baseline by more than 30% on average, while maintaining generation quality competitive with unwatermarked text and breaking the robustness-quality trade-off that limits prior methods. |
| title | SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness |
| topic | Cryptography and Security Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2605.25796 |