Fast Training Dataset Attribution via In-Context Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fotouhi, Milad, Bahadori, Mohammad Taha, Feyisetan, Oluwaseyi, Arabshahi, Payman, Heckerman, David
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913744124116992
author Fotouhi, Milad
Bahadori, Mohammad Taha
Feyisetan, Oluwaseyi
Arabshahi, Payman
Heckerman, David
author_facet Fotouhi, Milad
Bahadori, Mohammad Taha
Feyisetan, Oluwaseyi
Arabshahi, Payman
Heckerman, David
contents We investigate the use of in-context learning and prompt engineering to estimate the contributions of training data in the outputs of instruction-tuned large language models (LLMs). We propose two novel approaches: (1) a similarity-based approach that measures the difference between LLM outputs with and without provided context, and (2) a mixture distribution model approach that frames the problem of identifying contribution scores as a matrix factorization task. Our empirical comparison demonstrates that the mixture model approach is more robust to retrieval noise in in-context learning, providing a more reliable estimation of data contributions.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11852
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fast Training Dataset Attribution via In-Context Learning
Fotouhi, Milad
Bahadori, Mohammad Taha
Feyisetan, Oluwaseyi
Arabshahi, Payman
Heckerman, David
Computation and Language
Artificial Intelligence
Machine Learning
We investigate the use of in-context learning and prompt engineering to estimate the contributions of training data in the outputs of instruction-tuned large language models (LLMs). We propose two novel approaches: (1) a similarity-based approach that measures the difference between LLM outputs with and without provided context, and (2) a mixture distribution model approach that frames the problem of identifying contribution scores as a matrix factorization task. Our empirical comparison demonstrates that the mixture model approach is more robust to retrieval noise in in-context learning, providing a more reliable estimation of data contributions.
title Fast Training Dataset Attribution via In-Context Learning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2408.11852