How Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical Guidelines

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Deng, Junwei, Hu, Pingbang, Jin, Suliang, Lu, Hao, Wang, Jiachen T., Zhang, Shichang, Ma, Jiaqi W.
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914578998231040
author Deng, Junwei
Hu, Pingbang
Jin, Suliang
Lu, Hao
Wang, Jiachen T.
Zhang, Shichang
Ma, Jiaqi W.
author_facet Deng, Junwei
Hu, Pingbang
Jin, Suliang
Lu, Hao
Wang, Jiachen T.
Zhang, Shichang
Ma, Jiaqi W.
contents Trajectory-based data attribution methods estimate the influence of training samples on model predictions by unrolling the training trajectory. They are widely used in applications such as data selection, data valuation, and model diagnosis, but there is a lack of comprehensive error analysis of these methods, raising concerns about method faithfulness and hindering reliable deployment. In this work, we provide the first systematic analysis of error sources in trajectory-based data attribution, together with concrete remedies to mitigate them and practical guidelines for downstream use. We organize the total error into three categories, config-level, algorithm-level, and system-level. We make three contributions. First, we identify optimizer mismatch as the dominant config-level error: existing methods derive their attribution under the assumption of SGD, even for models trained with the modern de facto optimizer AdamW. We propose AdamW-influence to fully account for AdamW's optimization dynamics, yielding improvements from 10% to over 300% in Spearman correlation between estimated and ground-truth influence across four settings spanning MLP, CNN, GPT-2, and Llama 3.2-1B. Second, we isolate the remaining algorithm-level error arising from the first-order Taylor approximation, identify the learning rate and trajectory length as factors governing the error magnitude, and derive a closed-form error proxy that can be evaluated along the original trajectory without retraining. Third, we translate these insights into practical guidelines for data selection by unifying offline and online strategies under a K-step look-ahead framework. Under this framework, online selection with a short horizon often matches or exceeds offline, and the optimal horizon can be tuned jointly with the learning rate. Together, these results turn the framework into an actionable selection recipe for practitioners.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18814
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical Guidelines
Deng, Junwei
Hu, Pingbang
Jin, Suliang
Lu, Hao
Wang, Jiachen T.
Zhang, Shichang
Ma, Jiaqi W.
Machine Learning
Trajectory-based data attribution methods estimate the influence of training samples on model predictions by unrolling the training trajectory. They are widely used in applications such as data selection, data valuation, and model diagnosis, but there is a lack of comprehensive error analysis of these methods, raising concerns about method faithfulness and hindering reliable deployment. In this work, we provide the first systematic analysis of error sources in trajectory-based data attribution, together with concrete remedies to mitigate them and practical guidelines for downstream use. We organize the total error into three categories, config-level, algorithm-level, and system-level. We make three contributions. First, we identify optimizer mismatch as the dominant config-level error: existing methods derive their attribution under the assumption of SGD, even for models trained with the modern de facto optimizer AdamW. We propose AdamW-influence to fully account for AdamW's optimization dynamics, yielding improvements from 10% to over 300% in Spearman correlation between estimated and ground-truth influence across four settings spanning MLP, CNN, GPT-2, and Llama 3.2-1B. Second, we isolate the remaining algorithm-level error arising from the first-order Taylor approximation, identify the learning rate and trajectory length as factors governing the error magnitude, and derive a closed-form error proxy that can be evaluated along the original trajectory without retraining. Third, we translate these insights into practical guidelines for data selection by unifying offline and online strategies under a K-step look-ahead framework. Under this framework, online selection with a short horizon often matches or exceeds offline, and the optimal horizon can be tuned jointly with the learning rate. Together, these results turn the framework into an actionable selection recipe for practitioners.
title How Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical Guidelines
topic Machine Learning
url https://arxiv.org/abs/2605.18814