CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yilun, Tao, Shimin, Zhao, Xiaofeng, Zhu, Ming, Ma, Wenbing, Zhu, Junhao, Su, Chang, Hou, Yutai, Zhang, Miao, Zhang, Min, Ma, Hongxia, Zhang, Li, Yang, Hao, Jiang, Yanfei
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911806353571840
author Liu, Yilun
Tao, Shimin
Zhao, Xiaofeng
Zhu, Ming
Ma, Wenbing
Zhu, Junhao
Su, Chang
Hou, Yutai
Zhang, Miao
Zhang, Min
Ma, Hongxia
Zhang, Li
Yang, Hao
Jiang, Yanfei
author_facet Liu, Yilun
Tao, Shimin
Zhao, Xiaofeng
Zhu, Ming
Ma, Wenbing
Zhu, Junhao
Su, Chang
Hou, Yutai
Zhang, Miao
Zhang, Min
Ma, Hongxia
Zhang, Li
Yang, Hao
Jiang, Yanfei
contents Instruction tuning is crucial for enabling Language Learning Models (LLMs) in responding to human instructions. The quality of instruction pairs used for tuning greatly affects the performance of LLMs. However, the manual creation of high-quality instruction datasets is costly, leading to the adoption of automatic generation of instruction pairs by LLMs as a popular alternative. To ensure the high quality of LLM-generated instruction datasets, several approaches have been proposed. Nevertheless, existing methods either compromise dataset integrity by filtering a large proportion of samples, or are unsuitable for industrial applications. In this paper, instead of discarding low-quality samples, we propose CoachLM, a novel approach to enhance the quality of instruction datasets through automatic revisions on samples in the dataset. CoachLM is trained from the samples revised by human experts and significantly increases the proportion of high-quality samples in the dataset from 17.7% to 78.9%. The effectiveness of CoachLM is further assessed on various real-world instruction test sets. The results show that CoachLM improves the instruction-following capabilities of the instruction-tuned LLM by an average of 29.9%, which even surpasses larger LLMs with nearly twice the number of parameters. Furthermore, CoachLM is successfully deployed in a data management system for LLMs at Huawei, resulting in an efficiency improvement of up to 20% in the cleaning of 40k real-world instruction pairs. We release various assets of CoachLM, including the training data, code and test set (https://github.com/lunyiliu/CoachLM).
format Preprint
id arxiv_https___arxiv_org_abs_2311_13246
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction Tuning
Liu, Yilun
Tao, Shimin
Zhao, Xiaofeng
Zhu, Ming
Ma, Wenbing
Zhu, Junhao
Su, Chang
Hou, Yutai
Zhang, Miao
Zhang, Min
Ma, Hongxia
Zhang, Li
Yang, Hao
Jiang, Yanfei
Computation and Language
Instruction tuning is crucial for enabling Language Learning Models (LLMs) in responding to human instructions. The quality of instruction pairs used for tuning greatly affects the performance of LLMs. However, the manual creation of high-quality instruction datasets is costly, leading to the adoption of automatic generation of instruction pairs by LLMs as a popular alternative. To ensure the high quality of LLM-generated instruction datasets, several approaches have been proposed. Nevertheless, existing methods either compromise dataset integrity by filtering a large proportion of samples, or are unsuitable for industrial applications. In this paper, instead of discarding low-quality samples, we propose CoachLM, a novel approach to enhance the quality of instruction datasets through automatic revisions on samples in the dataset. CoachLM is trained from the samples revised by human experts and significantly increases the proportion of high-quality samples in the dataset from 17.7% to 78.9%. The effectiveness of CoachLM is further assessed on various real-world instruction test sets. The results show that CoachLM improves the instruction-following capabilities of the instruction-tuned LLM by an average of 29.9%, which even surpasses larger LLMs with nearly twice the number of parameters. Furthermore, CoachLM is successfully deployed in a data management system for LLMs at Huawei, resulting in an efficiency improvement of up to 20% in the cleaning of 40k real-world instruction pairs. We release various assets of CoachLM, including the training data, code and test set (https://github.com/lunyiliu/CoachLM).
title CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction Tuning
topic Computation and Language
url https://arxiv.org/abs/2311.13246