Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Shaobo, Jin, Xiangqi, Wang, Ziming, Wang, Jize, Zhang, Jiajun, Li, Kaixin, Wen, Zichen, Li, Zhong, He, Conghui, Hu, Xuming, Zhang, Linfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915316917862400
author Wang, Shaobo
Jin, Xiangqi
Wang, Ziming
Wang, Jize
Zhang, Jiajun
Li, Kaixin
Wen, Zichen
Li, Zhong
He, Conghui
Hu, Xuming
Zhang, Linfeng
author_facet Wang, Shaobo
Jin, Xiangqi
Wang, Ziming
Wang, Jize
Zhang, Jiajun
Li, Kaixin
Wen, Zichen
Li, Zhong
He, Conghui
Hu, Xuming
Zhang, Linfeng
contents Fine-tuning large language models (LLMs) on task-specific data is essential for their effective deployment. As dataset sizes grow, efficiently selecting optimal subsets for training becomes crucial to balancing performance and computational costs. Traditional data selection methods often require fine-tuning a scoring model on the target dataset, which is time-consuming and resource-intensive, or rely on heuristics that fail to fully leverage the model's predictive capabilities. To address these challenges, we propose Data Whisperer, an efficient, training-free, attention-based method that leverages few-shot in-context learning with the model to be fine-tuned. Comprehensive evaluations were conducted on both raw and synthetic datasets across diverse tasks and models. Notably, Data Whisperer achieves superior performance compared to the full GSM8K dataset on the Llama-3-8B-Instruct model, using just 10% of the data, and outperforms existing methods with a 3.1-point improvement and a 7.4$\times$ speedup. The code is available at https://github.com/gszfwsb/Data-Whisperer.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12212
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context Learning
Wang, Shaobo
Jin, Xiangqi
Wang, Ziming
Wang, Jize
Zhang, Jiajun
Li, Kaixin
Wen, Zichen
Li, Zhong
He, Conghui
Hu, Xuming
Zhang, Linfeng
Computation and Language
Fine-tuning large language models (LLMs) on task-specific data is essential for their effective deployment. As dataset sizes grow, efficiently selecting optimal subsets for training becomes crucial to balancing performance and computational costs. Traditional data selection methods often require fine-tuning a scoring model on the target dataset, which is time-consuming and resource-intensive, or rely on heuristics that fail to fully leverage the model's predictive capabilities. To address these challenges, we propose Data Whisperer, an efficient, training-free, attention-based method that leverages few-shot in-context learning with the model to be fine-tuned. Comprehensive evaluations were conducted on both raw and synthetic datasets across diverse tasks and models. Notably, Data Whisperer achieves superior performance compared to the full GSM8K dataset on the Llama-3-8B-Instruct model, using just 10% of the data, and outperforms existing methods with a 3.1-point improvement and a 7.4$\times$ speedup. The code is available at https://github.com/gszfwsb/Data-Whisperer.
title Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context Learning
topic Computation and Language
url https://arxiv.org/abs/2505.12212