De-mark: Watermark Removal in Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Ruibo, Wu, Yihan, Guo, Junfeng, Huang, Heng
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913923288006656
author Chen, Ruibo
Wu, Yihan
Guo, Junfeng
Huang, Heng
author_facet Chen, Ruibo
Wu, Yihan
Guo, Junfeng
Huang, Heng
contents Watermarking techniques offer a promising way to identify machine-generated content via embedding covert information into the contents generated from language models (LMs). However, the robustness of the watermarking schemes has not been well explored. In this paper, we present De-mark, an advanced framework designed to remove n-gram-based watermarks effectively. Our method utilizes a novel querying strategy, termed random selection probing, which aids in assessing the strength of the watermark and identifying the red-green list within the n-gram watermark. Experiments on popular LMs, such as Llama3 and ChatGPT, demonstrate the efficiency and effectiveness of De-mark in watermark removal and exploitation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13808
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle De-mark: Watermark Removal in Large Language Models
Chen, Ruibo
Wu, Yihan
Guo, Junfeng
Huang, Heng
Computation and Language
Watermarking techniques offer a promising way to identify machine-generated content via embedding covert information into the contents generated from language models (LMs). However, the robustness of the watermarking schemes has not been well explored. In this paper, we present De-mark, an advanced framework designed to remove n-gram-based watermarks effectively. Our method utilizes a novel querying strategy, termed random selection probing, which aids in assessing the strength of the watermark and identifying the red-green list within the n-gram watermark. Experiments on popular LMs, such as Llama3 and ChatGPT, demonstrate the efficiency and effectiveness of De-mark in watermark removal and exploitation tasks.
title De-mark: Watermark Removal in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2410.13808