Copyright Detection in Large Language Models: An Ethical Approach to Generative AI Development

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Szczecina, David, Gaffori, Senan, Li, Edmond
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912973356793856
author Szczecina, David
Gaffori, Senan
Li, Edmond
author_facet Szczecina, David
Gaffori, Senan
Li, Edmond
contents The widespread use of Large Language Models (LLMs) raises critical concerns regarding the unauthorized inclusion of copyrighted content in training data. Existing detection frameworks, such as DE-COP, are computationally intensive, and largely inaccessible to independent creators. As legal scrutiny increases, there is a pressing need for a scalable, transparent, and user-friendly solution. This paper introduce an open-source copyright detection platform that enables content creators to verify whether their work was used in LLM training datasets. Our approach enhances existing methodologies by facilitating ease of use, improving similarity detection, optimizing dataset validation, and reducing computational overhead by 10-30% with efficient API calls. With an intuitive user interface and scalable backend, this framework contributes to increasing transparency in AI development and ethical compliance, facilitating the foundation for further research in responsible AI development and copyright enforcement.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20623
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Copyright Detection in Large Language Models: An Ethical Approach to Generative AI Development
Szczecina, David
Gaffori, Senan
Li, Edmond
Artificial Intelligence
68T50
I.2.7
The widespread use of Large Language Models (LLMs) raises critical concerns regarding the unauthorized inclusion of copyrighted content in training data. Existing detection frameworks, such as DE-COP, are computationally intensive, and largely inaccessible to independent creators. As legal scrutiny increases, there is a pressing need for a scalable, transparent, and user-friendly solution. This paper introduce an open-source copyright detection platform that enables content creators to verify whether their work was used in LLM training datasets. Our approach enhances existing methodologies by facilitating ease of use, improving similarity detection, optimizing dataset validation, and reducing computational overhead by 10-30% with efficient API calls. With an intuitive user interface and scalable backend, this framework contributes to increasing transparency in AI development and ethical compliance, facilitating the foundation for further research in responsible AI development and copyright enforcement.
title Copyright Detection in Large Language Models: An Ethical Approach to Generative AI Development
topic Artificial Intelligence
68T50
I.2.7
url https://arxiv.org/abs/2511.20623