A Bayesian Information-Theoretic Approach to Data Attribution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tailor, Dharmesh, Felicioni, Nicolò, Ciosek, Kamil
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911577957990400
author Tailor, Dharmesh
Felicioni, Nicolò
Ciosek, Kamil
author_facet Tailor, Dharmesh
Felicioni, Nicolò
Ciosek, Kamil
contents Training Data Attribution (TDA) seeks to trace model predictions back to influential training examples, enhancing interpretability and safety. We formulate TDA as a Bayesian information-theoretic problem: subsets are scored by the information loss they induce - the entropy increase at a query when removed. This criterion credits examples for resolving predictive uncertainty rather than label noise. To scale to modern networks, we approximate information loss using a Gaussian Process surrogate built from tangent features. We show this aligns with classical influence scores for single-example attribution while promoting diversity for subsets. For even larger-scale retrieval, we relax to an information-gain objective and add a variance correction for scalable attribution in vector databases. Experiments show competitive performance on counterfactual sensitivity, ground-truth retrieval and coreset selection, showing that our method scales to modern architectures while bridging principled measures with practice.
format Preprint
id arxiv_https___arxiv_org_abs_2604_03858
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Bayesian Information-Theoretic Approach to Data Attribution
Tailor, Dharmesh
Felicioni, Nicolò
Ciosek, Kamil
Machine Learning
Training Data Attribution (TDA) seeks to trace model predictions back to influential training examples, enhancing interpretability and safety. We formulate TDA as a Bayesian information-theoretic problem: subsets are scored by the information loss they induce - the entropy increase at a query when removed. This criterion credits examples for resolving predictive uncertainty rather than label noise. To scale to modern networks, we approximate information loss using a Gaussian Process surrogate built from tangent features. We show this aligns with classical influence scores for single-example attribution while promoting diversity for subsets. For even larger-scale retrieval, we relax to an information-gain objective and add a variance correction for scalable attribution in vector databases. Experiments show competitive performance on counterfactual sensitivity, ground-truth retrieval and coreset selection, showing that our method scales to modern architectures while bridging principled measures with practice.
title A Bayesian Information-Theoretic Approach to Data Attribution
topic Machine Learning
url https://arxiv.org/abs/2604.03858