Exploring Training Data Attribution under Limited Access Constraints

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Shiyuan, Deng, Junwei, Bae, Juhan, Ma, Jiaqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912588937297920
author Zhang, Shiyuan
Deng, Junwei
Bae, Juhan
Ma, Jiaqi
author_facet Zhang, Shiyuan
Deng, Junwei
Bae, Juhan
Ma, Jiaqi
contents Training data attribution (TDA) plays a critical role in understanding the influence of individual training data points on model predictions. Gradient-based TDA methods, popularized by \textit{influence function} for their superior performance, have been widely applied in data selection, data cleaning, data economics, and fact tracing. However, in real-world scenarios where commercial models are not publicly accessible and computational resources are limited, existing TDA methods are often constrained by their reliance on full model access and high computational costs. This poses significant challenges to the broader adoption of TDA in practical applications. In this work, we present a systematic study of TDA methods under various access and resource constraints. We investigate the feasibility of performing TDA under varying levels of access constraints by leveraging appropriately designed solutions such as proxy models. Besides, we demonstrate that attribution scores obtained from models without prior training on the target dataset remain informative across a range of tasks, which is useful for scenarios where computational resources are limited. Our findings provide practical guidance for deploying TDA in real-world environments, aiming to improve feasibility and efficiency under limited access.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12581
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Training Data Attribution under Limited Access Constraints
Zhang, Shiyuan
Deng, Junwei
Bae, Juhan
Ma, Jiaqi
Machine Learning
Training data attribution (TDA) plays a critical role in understanding the influence of individual training data points on model predictions. Gradient-based TDA methods, popularized by \textit{influence function} for their superior performance, have been widely applied in data selection, data cleaning, data economics, and fact tracing. However, in real-world scenarios where commercial models are not publicly accessible and computational resources are limited, existing TDA methods are often constrained by their reliance on full model access and high computational costs. This poses significant challenges to the broader adoption of TDA in practical applications. In this work, we present a systematic study of TDA methods under various access and resource constraints. We investigate the feasibility of performing TDA under varying levels of access constraints by leveraging appropriately designed solutions such as proxy models. Besides, we demonstrate that attribution scores obtained from models without prior training on the target dataset remain informative across a range of tasks, which is useful for scenarios where computational resources are limited. Our findings provide practical guidance for deploying TDA in real-world environments, aiming to improve feasibility and efficiency under limited access.
title Exploring Training Data Attribution under Limited Access Constraints
topic Machine Learning
url https://arxiv.org/abs/2509.12581