Nine Ways to Break Copyright Law and Why Our LLM Won't: A Fair Use Aligned Generation Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sharma, Aakash Sen, Sanyal, Debdeep, Srivastava, Priyansh, H., Sundar Atreya, Karande, Shirish, Kankanhalli, Mohan, Mandal, Murari
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910974941855744
author Sharma, Aakash Sen
Sanyal, Debdeep
Srivastava, Priyansh
H., Sundar Atreya
Karande, Shirish
Kankanhalli, Mohan
Mandal, Murari
author_facet Sharma, Aakash Sen
Sanyal, Debdeep
Srivastava, Priyansh
H., Sundar Atreya
Karande, Shirish
Kankanhalli, Mohan
Mandal, Murari
contents Large language models (LLMs) commonly risk copyright infringement by reproducing protected content verbatim or with insufficient transformative modifications, posing significant ethical, legal, and practical concerns. Current inference-time safeguards predominantly rely on restrictive refusal-based filters, often compromising the practical utility of these models. To address this, we collaborated closely with intellectual property experts to develop FUA-LLM (Fair Use Aligned Language Models), a legally-grounded framework explicitly designed to align LLM outputs with fair-use doctrine. Central to our method is FairUseDB, a carefully constructed dataset containing 18,000 expert-validated examples covering nine realistic infringement scenarios. Leveraging this dataset, we apply Direct Preference Optimization (DPO) to fine-tune open-source LLMs, encouraging them to produce legally compliant and practically useful alternatives rather than resorting to blunt refusal. Recognizing the shortcomings of traditional evaluation metrics, we propose new measures: Weighted Penalty Utility and Compliance Aware Harmonic Mean (CAH) to balance infringement risk against response utility. Extensive quantitative experiments coupled with expert evaluations confirm that FUA-LLM substantially reduces problematic outputs (up to 20\%) compared to state-of-the-art approaches, while preserving real-world usability.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23788
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Nine Ways to Break Copyright Law and Why Our LLM Won't: A Fair Use Aligned Generation Framework
Sharma, Aakash Sen
Sanyal, Debdeep
Srivastava, Priyansh
H., Sundar Atreya
Karande, Shirish
Kankanhalli, Mohan
Mandal, Murari
Computation and Language
Artificial Intelligence
Large language models (LLMs) commonly risk copyright infringement by reproducing protected content verbatim or with insufficient transformative modifications, posing significant ethical, legal, and practical concerns. Current inference-time safeguards predominantly rely on restrictive refusal-based filters, often compromising the practical utility of these models. To address this, we collaborated closely with intellectual property experts to develop FUA-LLM (Fair Use Aligned Language Models), a legally-grounded framework explicitly designed to align LLM outputs with fair-use doctrine. Central to our method is FairUseDB, a carefully constructed dataset containing 18,000 expert-validated examples covering nine realistic infringement scenarios. Leveraging this dataset, we apply Direct Preference Optimization (DPO) to fine-tune open-source LLMs, encouraging them to produce legally compliant and practically useful alternatives rather than resorting to blunt refusal. Recognizing the shortcomings of traditional evaluation metrics, we propose new measures: Weighted Penalty Utility and Compliance Aware Harmonic Mean (CAH) to balance infringement risk against response utility. Extensive quantitative experiments coupled with expert evaluations confirm that FUA-LLM substantially reduces problematic outputs (up to 20\%) compared to state-of-the-art approaches, while preserving real-world usability.
title Nine Ways to Break Copyright Law and Why Our LLM Won't: A Fair Use Aligned Generation Framework
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.23788