Token-wise Influential Training Data Retrieval for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Huawei, Long, Jikai, Xu, Zhaozhuo, Zhao, Weijie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913559881973760
author Lin, Huawei
Long, Jikai
Xu, Zhaozhuo
Zhao, Weijie
author_facet Lin, Huawei
Long, Jikai
Xu, Zhaozhuo
Zhao, Weijie
contents Given a Large Language Model (LLM) generation, how can we identify which training data led to this generation? In this paper, we proposed RapidIn, a scalable framework adapting to LLMs for estimating the influence of each training data. The proposed framework consists of two stages: caching and retrieval. First, we compress the gradient vectors by over 200,000x, allowing them to be cached on disk or in GPU/CPU memory. Then, given a generation, RapidIn efficiently traverses the cached gradients to estimate the influence within minutes, achieving over a 6,326x speedup. Moreover, RapidIn supports multi-GPU parallelization to substantially accelerate caching and retrieval. Our empirical result confirms the efficiency and effectiveness of RapidIn.
format Preprint
id arxiv_https___arxiv_org_abs_2405_11724
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Token-wise Influential Training Data Retrieval for Large Language Models
Lin, Huawei
Long, Jikai
Xu, Zhaozhuo
Zhao, Weijie
Computation and Language
Artificial Intelligence
Cryptography and Security
Information Retrieval
Given a Large Language Model (LLM) generation, how can we identify which training data led to this generation? In this paper, we proposed RapidIn, a scalable framework adapting to LLMs for estimating the influence of each training data. The proposed framework consists of two stages: caching and retrieval. First, we compress the gradient vectors by over 200,000x, allowing them to be cached on disk or in GPU/CPU memory. Then, given a generation, RapidIn efficiently traverses the cached gradients to estimate the influence within minutes, achieving over a 6,326x speedup. Moreover, RapidIn supports multi-GPU parallelization to substantially accelerate caching and retrieval. Our empirical result confirms the efficiency and effectiveness of RapidIn.
title Token-wise Influential Training Data Retrieval for Large Language Models
topic Computation and Language
Artificial Intelligence
Cryptography and Security
Information Retrieval
url https://arxiv.org/abs/2405.11724