VOW: Verifiable and Oblivious Watermark Detection for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luan, Xiaokun, Zhang, Yihao, Su, Pengcheng, Lei, Feiran, Sun, Meng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915970630549504
author Luan, Xiaokun
Zhang, Yihao
Su, Pengcheng
Lei, Feiran
Sun, Meng
author_facet Luan, Xiaokun
Zhang, Yihao
Su, Pengcheng
Lei, Feiran
Sun, Meng
contents Large Language Model (LLM) watermarking is crucial for establishing the provenance of machine-generated text, but most existing methods rely on a centralized trust model. This model forces users to reveal potentially sensitive text to a provider for detection and offers no way to verify the integrity of the result. While asymmetric schemes have been proposed to address these issues, they are either impractical for short texts or lack formal guarantees linking watermark insertion and detection. We propose VOW, a new protocol that achieves both privacy-preserving and cryptographically verifiable watermark detection with high efficiency. Our approach formulates detection as a secure two-party computation problem, instantiating the watermark's core logic with a Verifiable Oblivious Pseudorandom Function (VOPRF). This allows the user and provider to perform detection without the user's text being revealed, while the provider's result is verifiable. Our comprehensive evaluation shows that VOW is practical for short texts and provides a crucial reassessment of watermark robustness against modern paraphrasing attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27666
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VOW: Verifiable and Oblivious Watermark Detection for Large Language Models
Luan, Xiaokun
Zhang, Yihao
Su, Pengcheng
Lei, Feiran
Sun, Meng
Cryptography and Security
Large Language Model (LLM) watermarking is crucial for establishing the provenance of machine-generated text, but most existing methods rely on a centralized trust model. This model forces users to reveal potentially sensitive text to a provider for detection and offers no way to verify the integrity of the result. While asymmetric schemes have been proposed to address these issues, they are either impractical for short texts or lack formal guarantees linking watermark insertion and detection. We propose VOW, a new protocol that achieves both privacy-preserving and cryptographically verifiable watermark detection with high efficiency. Our approach formulates detection as a secure two-party computation problem, instantiating the watermark's core logic with a Verifiable Oblivious Pseudorandom Function (VOPRF). This allows the user and provider to perform detection without the user's text being revealed, while the provider's result is verifiable. Our comprehensive evaluation shows that VOW is practical for short texts and provides a crucial reassessment of watermark robustness against modern paraphrasing attacks.
title VOW: Verifiable and Oblivious Watermark Detection for Large Language Models
topic Cryptography and Security
url https://arxiv.org/abs/2604.27666