Fast Training Dataset Attribution via In-Context Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913744124116992 |
|---|---|
| author | Fotouhi, Milad Bahadori, Mohammad Taha Feyisetan, Oluwaseyi Arabshahi, Payman Heckerman, David |
| author_facet | Fotouhi, Milad Bahadori, Mohammad Taha Feyisetan, Oluwaseyi Arabshahi, Payman Heckerman, David |
| contents | We investigate the use of in-context learning and prompt engineering to estimate the contributions of training data in the outputs of instruction-tuned large language models (LLMs). We propose two novel approaches: (1) a similarity-based approach that measures the difference between LLM outputs with and without provided context, and (2) a mixture distribution model approach that frames the problem of identifying contribution scores as a matrix factorization task. Our empirical comparison demonstrates that the mixture model approach is more robust to retrieval noise in in-context learning, providing a more reliable estimation of data contributions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_11852 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Fast Training Dataset Attribution via In-Context Learning Fotouhi, Milad Bahadori, Mohammad Taha Feyisetan, Oluwaseyi Arabshahi, Payman Heckerman, David Computation and Language Artificial Intelligence Machine Learning We investigate the use of in-context learning and prompt engineering to estimate the contributions of training data in the outputs of instruction-tuned large language models (LLMs). We propose two novel approaches: (1) a similarity-based approach that measures the difference between LLM outputs with and without provided context, and (2) a mixture distribution model approach that frames the problem of identifying contribution scores as a matrix factorization task. Our empirical comparison demonstrates that the mixture model approach is more robust to retrieval noise in in-context learning, providing a more reliable estimation of data contributions. |
| title | Fast Training Dataset Attribution via In-Context Learning |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2408.11852 |