Token-wise Influential Training Data Retrieval for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913559881973760 |
|---|---|
| author | Lin, Huawei Long, Jikai Xu, Zhaozhuo Zhao, Weijie |
| author_facet | Lin, Huawei Long, Jikai Xu, Zhaozhuo Zhao, Weijie |
| contents | Given a Large Language Model (LLM) generation, how can we identify which training data led to this generation? In this paper, we proposed RapidIn, a scalable framework adapting to LLMs for estimating the influence of each training data. The proposed framework consists of two stages: caching and retrieval. First, we compress the gradient vectors by over 200,000x, allowing them to be cached on disk or in GPU/CPU memory. Then, given a generation, RapidIn efficiently traverses the cached gradients to estimate the influence within minutes, achieving over a 6,326x speedup. Moreover, RapidIn supports multi-GPU parallelization to substantially accelerate caching and retrieval. Our empirical result confirms the efficiency and effectiveness of RapidIn. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_11724 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Token-wise Influential Training Data Retrieval for Large Language Models Lin, Huawei Long, Jikai Xu, Zhaozhuo Zhao, Weijie Computation and Language Artificial Intelligence Cryptography and Security Information Retrieval Given a Large Language Model (LLM) generation, how can we identify which training data led to this generation? In this paper, we proposed RapidIn, a scalable framework adapting to LLMs for estimating the influence of each training data. The proposed framework consists of two stages: caching and retrieval. First, we compress the gradient vectors by over 200,000x, allowing them to be cached on disk or in GPU/CPU memory. Then, given a generation, RapidIn efficiently traverses the cached gradients to estimate the influence within minutes, achieving over a 6,326x speedup. Moreover, RapidIn supports multi-GPU parallelization to substantially accelerate caching and retrieval. Our empirical result confirms the efficiency and effectiveness of RapidIn. |
| title | Token-wise Influential Training Data Retrieval for Large Language Models |
| topic | Computation and Language Artificial Intelligence Cryptography and Security Information Retrieval |
| url | https://arxiv.org/abs/2405.11724 |