Watermarking Text Data on Large Language Models for Dataset Copyright

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yixin, Hu, Hongsheng, Chen, Xun, Zhang, Xuyun, Sun, Lichao
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909252522606592
author Liu, Yixin
Hu, Hongsheng
Chen, Xun
Zhang, Xuyun
Sun, Lichao
author_facet Liu, Yixin
Hu, Hongsheng
Chen, Xun
Zhang, Xuyun
Sun, Lichao
contents Substantial research works have shown that deep models, e.g., pre-trained models, on the large corpus can learn universal language representations, which are beneficial for downstream NLP tasks. However, these powerful models are also vulnerable to various privacy attacks, while much sensitive information exists in the training dataset. The attacker can easily steal sensitive information from public models, e.g., individuals' email addresses and phone numbers. In an attempt to address these issues, particularly the unauthorized use of private data, we introduce a novel watermarking technique via a backdoor-based membership inference approach named TextMarker, which can safeguard diverse forms of private information embedded in the training text data. Specifically, TextMarker only requires data owners to mark a small number of samples for data copyright protection under the black-box access assumption to the target model. Through extensive evaluation, we demonstrate the effectiveness of TextMarker on various real-world datasets, e.g., marking only 0.1% of the training dataset is practically sufficient for effective membership inference with negligible effect on model utility. We also discuss potential countermeasures and show that TextMarker is stealthy enough to bypass them.
format Preprint
id arxiv_https___arxiv_org_abs_2305_13257
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Watermarking Text Data on Large Language Models for Dataset Copyright
Liu, Yixin
Hu, Hongsheng
Chen, Xun
Zhang, Xuyun
Sun, Lichao
Cryptography and Security
Substantial research works have shown that deep models, e.g., pre-trained models, on the large corpus can learn universal language representations, which are beneficial for downstream NLP tasks. However, these powerful models are also vulnerable to various privacy attacks, while much sensitive information exists in the training dataset. The attacker can easily steal sensitive information from public models, e.g., individuals' email addresses and phone numbers. In an attempt to address these issues, particularly the unauthorized use of private data, we introduce a novel watermarking technique via a backdoor-based membership inference approach named TextMarker, which can safeguard diverse forms of private information embedded in the training text data. Specifically, TextMarker only requires data owners to mark a small number of samples for data copyright protection under the black-box access assumption to the target model. Through extensive evaluation, we demonstrate the effectiveness of TextMarker on various real-world datasets, e.g., marking only 0.1% of the training dataset is practically sufficient for effective membership inference with negligible effect on model utility. We also discuss potential countermeasures and show that TextMarker is stealthy enough to bypass them.
title Watermarking Text Data on Large Language Models for Dataset Copyright
topic Cryptography and Security
url https://arxiv.org/abs/2305.13257