Generalized Group Data Attribution

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ley, Dan, Srinivas, Suraj, Zhang, Shichang, Rusak, Gili, Lakkaraju, Himabindu
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912078243037184
author Ley, Dan
Srinivas, Suraj
Zhang, Shichang
Rusak, Gili
Lakkaraju, Himabindu
author_facet Ley, Dan
Srinivas, Suraj
Zhang, Shichang
Rusak, Gili
Lakkaraju, Himabindu
contents Data Attribution (DA) methods quantify the influence of individual training data points on model outputs and have broad applications such as explainability, data selection, and noisy label identification. However, existing DA methods are often computationally intensive, limiting their applicability to large-scale machine learning models. To address this challenge, we introduce the Generalized Group Data Attribution (GGDA) framework, which computationally simplifies DA by attributing to groups of training points instead of individual ones. GGDA is a general framework that subsumes existing attribution methods and can be applied to new DA techniques as they emerge. It allows users to optimize the trade-off between efficiency and fidelity based on their needs. Our empirical results demonstrate that GGDA applied to popular DA methods such as Influence Functions, TracIn, and TRAK results in upto 10x-50x speedups over standard DA methods while gracefully trading off attribution fidelity. For downstream applications such as dataset pruning and noisy label identification, we demonstrate that GGDA significantly improves computational efficiency and maintains effectiveness, enabling practical applications in large-scale machine learning scenarios that were previously infeasible.
format Preprint
id arxiv_https___arxiv_org_abs_2410_09940
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generalized Group Data Attribution
Ley, Dan
Srinivas, Suraj
Zhang, Shichang
Rusak, Gili
Lakkaraju, Himabindu
Machine Learning
Artificial Intelligence
Data Attribution (DA) methods quantify the influence of individual training data points on model outputs and have broad applications such as explainability, data selection, and noisy label identification. However, existing DA methods are often computationally intensive, limiting their applicability to large-scale machine learning models. To address this challenge, we introduce the Generalized Group Data Attribution (GGDA) framework, which computationally simplifies DA by attributing to groups of training points instead of individual ones. GGDA is a general framework that subsumes existing attribution methods and can be applied to new DA techniques as they emerge. It allows users to optimize the trade-off between efficiency and fidelity based on their needs. Our empirical results demonstrate that GGDA applied to popular DA methods such as Influence Functions, TracIn, and TRAK results in upto 10x-50x speedups over standard DA methods while gracefully trading off attribution fidelity. For downstream applications such as dataset pruning and noisy label identification, we demonstrate that GGDA significantly improves computational efficiency and maintains effectiveness, enabling practical applications in large-scale machine learning scenarios that were previously infeasible.
title Generalized Group Data Attribution
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.09940