Enhancing Training Data Attribution with Representational Optimization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sun, Weiwei, Liu, Haokun, Kandpal, Nikhil, Raffel, Colin, Yang, Yiming
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914171068612608
author Sun, Weiwei
Liu, Haokun
Kandpal, Nikhil
Raffel, Colin
Yang, Yiming
author_facet Sun, Weiwei
Liu, Haokun
Kandpal, Nikhil
Raffel, Colin
Yang, Yiming
contents Training data attribution (TDA) methods aim to measure how training data impacts a model's predictions. While gradient-based attribution methods, such as influence functions, offer theoretical grounding, their computational costs make them impractical for large-scale applications. Representation-based approaches are far more scalable, but typically rely on heuristic embeddings that are not optimized for attribution, limiting their fidelity. To address these challenges, we propose AirRep, a scalable, representation-based approach that closes this gap by learning task-specific and model-aligned representations optimized explicitly for TDA. AirRep introduces two key innovations: a trainable encoder tuned for attribution quality, and an attention-based pooling mechanism that enables accurate estimation of group-wise influence. We train AirRep using a ranking objective over automatically constructed training subsets labeled by their empirical effect on target predictions. Experiments on instruction-tuned LLMs demonstrate that AirRep achieves performance on par with state-of-the-art gradient-based approaches while being nearly two orders of magnitude more efficient at inference time. Further analysis highlights its robustness and generalization across tasks and models. Our code is available at https://github.com/sunnweiwei/AirRep
format Preprint
id arxiv_https___arxiv_org_abs_2505_18513
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Training Data Attribution with Representational Optimization
Sun, Weiwei
Liu, Haokun
Kandpal, Nikhil
Raffel, Colin
Yang, Yiming
Machine Learning
Training data attribution (TDA) methods aim to measure how training data impacts a model's predictions. While gradient-based attribution methods, such as influence functions, offer theoretical grounding, their computational costs make them impractical for large-scale applications. Representation-based approaches are far more scalable, but typically rely on heuristic embeddings that are not optimized for attribution, limiting their fidelity. To address these challenges, we propose AirRep, a scalable, representation-based approach that closes this gap by learning task-specific and model-aligned representations optimized explicitly for TDA. AirRep introduces two key innovations: a trainable encoder tuned for attribution quality, and an attention-based pooling mechanism that enables accurate estimation of group-wise influence. We train AirRep using a ranking objective over automatically constructed training subsets labeled by their empirical effect on target predictions. Experiments on instruction-tuned LLMs demonstrate that AirRep achieves performance on par with state-of-the-art gradient-based approaches while being nearly two orders of magnitude more efficient at inference time. Further analysis highlights its robustness and generalization across tasks and models. Our code is available at https://github.com/sunnweiwei/AirRep
title Enhancing Training Data Attribution with Representational Optimization
topic Machine Learning
url https://arxiv.org/abs/2505.18513