SoK: Dataset Copyright Auditing in Machine Learning Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Du, Linkang, Zhou, Xuanru, Chen, Min, Zhang, Chusong, Su, Zhou, Cheng, Peng, Chen, Jiming, Zhang, Zhikun
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912390794182656
author Du, Linkang
Zhou, Xuanru
Chen, Min
Zhang, Chusong
Su, Zhou
Cheng, Peng
Chen, Jiming
Zhang, Zhikun
author_facet Du, Linkang
Zhou, Xuanru
Chen, Min
Zhang, Chusong
Su, Zhou
Cheng, Peng
Chen, Jiming
Zhang, Zhikun
contents As the implementation of machine learning (ML) systems becomes more widespread, especially with the introduction of larger ML models, we perceive a spring demand for massive data. However, it inevitably causes infringement and misuse problems with the data, such as using unauthorized online artworks or face images to train ML models. To address this problem, many efforts have been made to audit the copyright of the model training dataset. However, existing solutions vary in auditing assumptions and capabilities, making it difficult to compare their strengths and weaknesses. In addition, robustness evaluations usually consider only part of the ML pipeline and hardly reflect the performance of algorithms in real-world ML applications. Thus, it is essential to take a practical deployment perspective on the current dataset copyright auditing tools, examining their effectiveness and limitations. Concretely, we categorize dataset copyright auditing research into two prominent strands: intrusive methods and non-intrusive methods, depending on whether they require modifications to the original dataset. Then, we break down the intrusive methods into different watermark injection options and examine the non-intrusive methods using various fingerprints. To summarize our results, we offer detailed reference tables, highlight key points, and pinpoint unresolved issues in the current literature. By combining the pipeline in ML systems and analyzing previous studies, we highlight several future directions to make auditing tools more suitable for real-world copyright protection requirements.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16618
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SoK: Dataset Copyright Auditing in Machine Learning Systems
Du, Linkang
Zhou, Xuanru
Chen, Min
Zhang, Chusong
Su, Zhou
Cheng, Peng
Chen, Jiming
Zhang, Zhikun
Cryptography and Security
Machine Learning
As the implementation of machine learning (ML) systems becomes more widespread, especially with the introduction of larger ML models, we perceive a spring demand for massive data. However, it inevitably causes infringement and misuse problems with the data, such as using unauthorized online artworks or face images to train ML models. To address this problem, many efforts have been made to audit the copyright of the model training dataset. However, existing solutions vary in auditing assumptions and capabilities, making it difficult to compare their strengths and weaknesses. In addition, robustness evaluations usually consider only part of the ML pipeline and hardly reflect the performance of algorithms in real-world ML applications. Thus, it is essential to take a practical deployment perspective on the current dataset copyright auditing tools, examining their effectiveness and limitations. Concretely, we categorize dataset copyright auditing research into two prominent strands: intrusive methods and non-intrusive methods, depending on whether they require modifications to the original dataset. Then, we break down the intrusive methods into different watermark injection options and examine the non-intrusive methods using various fingerprints. To summarize our results, we offer detailed reference tables, highlight key points, and pinpoint unresolved issues in the current literature. By combining the pipeline in ML systems and analyzing previous studies, we highlight several future directions to make auditing tools more suitable for real-world copyright protection requirements.
title SoK: Dataset Copyright Auditing in Machine Learning Systems
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2410.16618