Randomization Techniques to Mitigate the Risk of Copyright Infringement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Wei-Ning, Kairouz, Peter, Oh, Sewoong, Xu, Zheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912000785776640
author Chen, Wei-Ning
Kairouz, Peter
Oh, Sewoong
Xu, Zheng
author_facet Chen, Wei-Ning
Kairouz, Peter
Oh, Sewoong
Xu, Zheng
contents In this paper, we investigate potential randomization approaches that can complement current practices of input-based methods (such as licensing data and prompt filtering) and output-based methods (such as recitation checker, license checker, and model-based similarity score) for copyright protection. This is motivated by the inherent ambiguity of the rules that determine substantial similarity in copyright precedents. Given that there is no quantifiable measure of substantial similarity that is agreed upon, complementary approaches can potentially further decrease liability. Similar randomized approaches, such as differential privacy, have been successful in mitigating privacy risks. This document focuses on the technical and research perspective on mitigating copyright violation and hence is not confidential. After investigating potential solutions and running numerical experiments, we concluded that using the notion of Near Access-Freeness (NAF) to measure the degree of substantial similarity is challenging, and the standard approach of training a Differentially Private (DP) model costs significantly when used to ensure NAF. Alternative approaches, such as retrieval models, might provide a more controllable scheme for mitigating substantial similarity.
format Preprint
id arxiv_https___arxiv_org_abs_2408_13278
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Randomization Techniques to Mitigate the Risk of Copyright Infringement
Chen, Wei-Ning
Kairouz, Peter
Oh, Sewoong
Xu, Zheng
Cryptography and Security
Machine Learning
In this paper, we investigate potential randomization approaches that can complement current practices of input-based methods (such as licensing data and prompt filtering) and output-based methods (such as recitation checker, license checker, and model-based similarity score) for copyright protection. This is motivated by the inherent ambiguity of the rules that determine substantial similarity in copyright precedents. Given that there is no quantifiable measure of substantial similarity that is agreed upon, complementary approaches can potentially further decrease liability. Similar randomized approaches, such as differential privacy, have been successful in mitigating privacy risks. This document focuses on the technical and research perspective on mitigating copyright violation and hence is not confidential. After investigating potential solutions and running numerical experiments, we concluded that using the notion of Near Access-Freeness (NAF) to measure the degree of substantial similarity is challenging, and the standard approach of training a Differentially Private (DP) model costs significantly when used to ensure NAF. Alternative approaches, such as retrieval models, might provide a more controllable scheme for mitigating substantial similarity.
title Randomization Techniques to Mitigate the Risk of Copyright Infringement
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2408.13278