Digger: Detecting Copyright Content Mis-usage in Large Language Model Training

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Haodong, Deng, Gelei, Liu, Yi, Wang, Kailong, Li, Yuekang, Zhang, Tianwei, Liu, Yang, Xu, Guoai, Xu, Guosheng, Wang, Haoyu
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929195864555520
author Li, Haodong
Deng, Gelei
Liu, Yi
Wang, Kailong
Li, Yuekang
Zhang, Tianwei
Liu, Yang
Xu, Guoai
Xu, Guosheng
Wang, Haoyu
author_facet Li, Haodong
Deng, Gelei
Liu, Yi
Wang, Kailong
Li, Yuekang
Zhang, Tianwei
Liu, Yang
Xu, Guoai
Xu, Guosheng
Wang, Haoyu
contents Pre-training, which utilizes extensive and varied datasets, is a critical factor in the success of Large Language Models (LLMs) across numerous applications. However, the detailed makeup of these datasets is often not disclosed, leading to concerns about data security and potential misuse. This is particularly relevant when copyrighted material, still under legal protection, is used inappropriately, either intentionally or unintentionally, infringing on the rights of the authors. In this paper, we introduce a detailed framework designed to detect and assess the presence of content from potentially copyrighted books within the training datasets of LLMs. This framework also provides a confidence estimation for the likelihood of each content sample's inclusion. To validate our approach, we conduct a series of simulated experiments, the results of which affirm the framework's effectiveness in identifying and addressing instances of content misuse in LLM training processes. Furthermore, we investigate the presence of recognizable quotes from famous literary works within these datasets. The outcomes of our study have significant implications for ensuring the ethical use of copyrighted materials in the development of LLMs, highlighting the need for more transparent and responsible data management practices in this field.
format Preprint
id arxiv_https___arxiv_org_abs_2401_00676
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Digger: Detecting Copyright Content Mis-usage in Large Language Model Training
Li, Haodong
Deng, Gelei
Liu, Yi
Wang, Kailong
Li, Yuekang
Zhang, Tianwei
Liu, Yang
Xu, Guoai
Xu, Guosheng
Wang, Haoyu
Cryptography and Security
Computation and Language
Machine Learning
Pre-training, which utilizes extensive and varied datasets, is a critical factor in the success of Large Language Models (LLMs) across numerous applications. However, the detailed makeup of these datasets is often not disclosed, leading to concerns about data security and potential misuse. This is particularly relevant when copyrighted material, still under legal protection, is used inappropriately, either intentionally or unintentionally, infringing on the rights of the authors. In this paper, we introduce a detailed framework designed to detect and assess the presence of content from potentially copyrighted books within the training datasets of LLMs. This framework also provides a confidence estimation for the likelihood of each content sample's inclusion. To validate our approach, we conduct a series of simulated experiments, the results of which affirm the framework's effectiveness in identifying and addressing instances of content misuse in LLM training processes. Furthermore, we investigate the presence of recognizable quotes from famous literary works within these datasets. The outcomes of our study have significant implications for ensuring the ethical use of copyrighted materials in the development of LLMs, highlighting the need for more transparent and responsible data management practices in this field.
title Digger: Detecting Copyright Content Mis-usage in Large Language Model Training
topic Cryptography and Security
Computation and Language
Machine Learning
url https://arxiv.org/abs/2401.00676