Beyond English: Unveiling Multilingual Bias in LLM Copyright Compliance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Yupeng, Zhang, Xiaoyu, Huang, Yixian, Xie, Qian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916646818414592
author Chen, Yupeng
Zhang, Xiaoyu
Huang, Yixian
Xie, Qian
author_facet Chen, Yupeng
Zhang, Xiaoyu
Huang, Yixian
Xie, Qian
contents Large Language Models (LLMs) have raised significant concerns regarding the fair use of copyright-protected content. While prior studies have examined the extent to which LLMs reproduce copyrighted materials, they have predominantly focused on English, neglecting multilingual dimensions of copyright protection. In this work, we investigate multilingual biases in LLM copyright protection by addressing two key questions: (1) Do LLMs exhibit bias in protecting copyrighted works across languages? (2) Is it easier to elicit copyrighted content using prompts in specific languages? To explore these questions, we construct a dataset of popular song lyrics in English, French, Chinese, and Korean and systematically probe seven LLMs using prompts in these languages. Our findings reveal significant imbalances in LLMs' handling of copyrighted content, both in terms of the language of the copyrighted material and the language of the prompt. These results highlight the need for further research and development of more robust, language-agnostic copyright protection mechanisms to ensure fair and consistent protection across languages.
format Preprint
id arxiv_https___arxiv_org_abs_2503_05713
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond English: Unveiling Multilingual Bias in LLM Copyright Compliance
Chen, Yupeng
Zhang, Xiaoyu
Huang, Yixian
Xie, Qian
Computers and Society
Computation and Language
Large Language Models (LLMs) have raised significant concerns regarding the fair use of copyright-protected content. While prior studies have examined the extent to which LLMs reproduce copyrighted materials, they have predominantly focused on English, neglecting multilingual dimensions of copyright protection. In this work, we investigate multilingual biases in LLM copyright protection by addressing two key questions: (1) Do LLMs exhibit bias in protecting copyrighted works across languages? (2) Is it easier to elicit copyrighted content using prompts in specific languages? To explore these questions, we construct a dataset of popular song lyrics in English, French, Chinese, and Korean and systematically probe seven LLMs using prompts in these languages. Our findings reveal significant imbalances in LLMs' handling of copyrighted content, both in terms of the language of the copyrighted material and the language of the prompt. These results highlight the need for further research and development of more robust, language-agnostic copyright protection mechanisms to ensure fair and consistent protection across languages.
title Beyond English: Unveiling Multilingual Bias in LLM Copyright Compliance
topic Computers and Society
Computation and Language
url https://arxiv.org/abs/2503.05713