From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Ming, Zhang, Yong, Li, Zhitao, Chen, Jiuhai, Chen, Lichang, Cheng, Ning, Wang, Jianzong, Zhou, Tianyi, Xiao, Jing
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911828716552192
author Li, Ming
Zhang, Yong
Li, Zhitao
Chen, Jiuhai
Chen, Lichang
Cheng, Ning
Wang, Jianzong
Zhou, Tianyi
Xiao, Jing
author_facet Li, Ming
Zhang, Yong
Li, Zhitao
Chen, Jiuhai
Chen, Lichang
Cheng, Ning
Wang, Jianzong
Zhou, Tianyi
Xiao, Jing
contents In the realm of Large Language Models (LLMs), the balance between instruction data quality and quantity is a focal point. Recognizing this, we introduce a self-guided methodology for LLMs to autonomously discern and select cherry samples from open-source datasets, effectively minimizing manual curation and potential cost for instruction tuning an LLM. Our key innovation, the Instruction-Following Difficulty (IFD) metric, emerges as a pivotal metric to identify discrepancies between a model's expected responses and its intrinsic generation capability. Through the application of IFD, cherry samples can be pinpointed, leading to a marked uptick in model training efficiency. Empirical validations on datasets like Alpaca and WizardLM underpin our findings; with a mere $10\%$ of original data input, our strategy showcases improved results. This synthesis of self-guided cherry-picking and the IFD metric signifies a transformative leap in the instruction tuning of LLMs, promising both efficiency and resource-conscious advancements. Codes, data, and models are available: https://github.com/tianyi-lab/Cherry_LLM
format Preprint
id arxiv_https___arxiv_org_abs_2308_12032
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
Li, Ming
Zhang, Yong
Li, Zhitao
Chen, Jiuhai
Chen, Lichang
Cheng, Ning
Wang, Jianzong
Zhou, Tianyi
Xiao, Jing
Computation and Language
In the realm of Large Language Models (LLMs), the balance between instruction data quality and quantity is a focal point. Recognizing this, we introduce a self-guided methodology for LLMs to autonomously discern and select cherry samples from open-source datasets, effectively minimizing manual curation and potential cost for instruction tuning an LLM. Our key innovation, the Instruction-Following Difficulty (IFD) metric, emerges as a pivotal metric to identify discrepancies between a model's expected responses and its intrinsic generation capability. Through the application of IFD, cherry samples can be pinpointed, leading to a marked uptick in model training efficiency. Empirical validations on datasets like Alpaca and WizardLM underpin our findings; with a mere $10\%$ of original data input, our strategy showcases improved results. This synthesis of self-guided cherry-picking and the IFD metric signifies a transformative leap in the instruction tuning of LLMs, promising both efficiency and resource-conscious advancements. Codes, data, and models are available: https://github.com/tianyi-lab/Cherry_LLM
title From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
topic Computation and Language
url https://arxiv.org/abs/2308.12032