Data Shapley in One Training Run

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jiachen T., Mittal, Prateek, Song, Dawn, Jia, Ruoxi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915332212391936
author Wang, Jiachen T.
Mittal, Prateek
Song, Dawn
Jia, Ruoxi
author_facet Wang, Jiachen T.
Mittal, Prateek
Song, Dawn
Jia, Ruoxi
contents Data Shapley provides a principled framework for attributing data's contribution within machine learning contexts. However, existing approaches require re-training models on different data subsets, which is computationally intensive, foreclosing their application to large-scale models. Furthermore, they produce the same attribution score for any models produced by running the learning algorithm, meaning they cannot perform targeted attribution towards a specific model obtained from a single run of the algorithm. This paper introduces In-Run Data Shapley, which addresses these limitations by offering scalable data attribution for a target model of interest. In its most efficient implementation, our technique incurs negligible additional runtime compared to standard model training. This dramatic efficiency improvement makes it possible to perform data attribution for the foundation model pretraining stage for the first time. We present several case studies that offer fresh insights into pretraining data's contribution and discuss their implications for copyright in generative AI and pretraining data curation.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11011
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data Shapley in One Training Run
Wang, Jiachen T.
Mittal, Prateek
Song, Dawn
Jia, Ruoxi
Machine Learning
Computation and Language
Data Shapley provides a principled framework for attributing data's contribution within machine learning contexts. However, existing approaches require re-training models on different data subsets, which is computationally intensive, foreclosing their application to large-scale models. Furthermore, they produce the same attribution score for any models produced by running the learning algorithm, meaning they cannot perform targeted attribution towards a specific model obtained from a single run of the algorithm. This paper introduces In-Run Data Shapley, which addresses these limitations by offering scalable data attribution for a target model of interest. In its most efficient implementation, our technique incurs negligible additional runtime compared to standard model training. This dramatic efficiency improvement makes it possible to perform data attribution for the foundation model pretraining stage for the first time. We present several case studies that offer fresh insights into pretraining data's contribution and discuss their implications for copyright in generative AI and pretraining data curation.
title Data Shapley in One Training Run
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2406.11011