RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Yixin, Dong, Qingxiu, Yao, Linli, Zhu, Fangwei, Sui, Zhifang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910949955338240
author Yang, Yixin
Dong, Qingxiu
Yao, Linli
Zhu, Fangwei
Sui, Zhifang
author_facet Yang, Yixin
Dong, Qingxiu
Yao, Linli
Zhu, Fangwei
Sui, Zhifang
contents Data selection for instruction tuning is crucial for improving the performance of large language models (LLMs) while reducing training costs. In this paper, we propose Refined Contribution Measurement with In-Context Learning (RICo), a novel gradient-free method that quantifies the fine-grained contribution of individual samples to both task-level and global-level model performance. RICo enables more accurate identification of high-contribution data, leading to better instruction tuning. We further introduce a lightweight selection paradigm trained on RICo scores, enabling scalable data selection with a strictly linear inference complexity. Extensive experiments on three LLMs across 12 benchmarks and 5 pairwise evaluation sets demonstrate the effectiveness of RICo. Remarkably, on LLaMA3.1-8B, models trained on 15% of RICo-selected data outperform full datasets by 5.42% points and exceed the best performance of widely used selection methods by 2.06% points. We further analyze high-contribution samples selected by RICo, which show both diverse tasks and appropriate difficulty levels, rather than just the hardest ones.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05327
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
Yang, Yixin
Dong, Qingxiu
Yao, Linli
Zhu, Fangwei
Sui, Zhifang
Computation and Language
Data selection for instruction tuning is crucial for improving the performance of large language models (LLMs) while reducing training costs. In this paper, we propose Refined Contribution Measurement with In-Context Learning (RICo), a novel gradient-free method that quantifies the fine-grained contribution of individual samples to both task-level and global-level model performance. RICo enables more accurate identification of high-contribution data, leading to better instruction tuning. We further introduce a lightweight selection paradigm trained on RICo scores, enabling scalable data selection with a strictly linear inference complexity. Extensive experiments on three LLMs across 12 benchmarks and 5 pairwise evaluation sets demonstrate the effectiveness of RICo. Remarkably, on LLaMA3.1-8B, models trained on 15% of RICo-selected data outperform full datasets by 5.42% points and exceed the best performance of widely used selection methods by 2.06% points. We further analyze high-contribution samples selected by RICo, which show both diverse tasks and appropriate difficulty levels, rather than just the hardest ones.
title RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
topic Computation and Language
url https://arxiv.org/abs/2505.05327