Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality Estimation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ge, Yuan, Liu, Yilun, Hu, Chi, Meng, Weibin, Tao, Shimin, Zhao, Xiaofeng, Ma, Hongxia, Zhang, Li, Chen, Boxing, Yang, Hao, Li, Bei, Xiao, Tong, Zhu, Jingbo
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915023754887168
author Ge, Yuan
Liu, Yilun
Hu, Chi
Meng, Weibin
Tao, Shimin
Zhao, Xiaofeng
Ma, Hongxia
Zhang, Li
Chen, Boxing
Yang, Hao
Li, Bei
Xiao, Tong
Zhu, Jingbo
author_facet Ge, Yuan
Liu, Yilun
Hu, Chi
Meng, Weibin
Tao, Shimin
Zhao, Xiaofeng
Ma, Hongxia
Zhang, Li
Chen, Boxing
Yang, Hao
Li, Bei
Xiao, Tong
Zhu, Jingbo
contents With contributions from the open-source community, a vast amount of instruction tuning (IT) data has emerged. Given the significant resource allocation required for training and evaluating models, it is advantageous to have an efficient method for selecting high-quality IT data. However, existing methods for instruction data selection have limitations such as relying on fragile external APIs, being affected by biases in GPT models, or reducing the diversity of the selected instruction dataset. In this paper, we propose an industrial-friendly, expert-aligned and diversity-preserved instruction data selection method: Clustering and Ranking (CaR). CaR employs a two-step process: first, it ranks instruction pairs using a high-accuracy (84.25%) scoring model aligned with expert preferences; second, it preserves dataset diversity through clustering. In our experiment, CaR efficiently selected a mere 1.96% of Alpaca's IT data, yet the resulting AlpaCaR model surpassed Alpaca's performance by an average of 32.1% in GPT-4 evaluations. Moreover, we find that data selecting is a consistent paradigm whether the pre-trained model is more capable or the model parameters scaling up. Our approach employs compact models with 550M parameters and incurs just 11.2% of the financial outlay of current methods, enhancing its industrial deployability.
format Preprint
id arxiv_https___arxiv_org_abs_2402_18191
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality Estimation
Ge, Yuan
Liu, Yilun
Hu, Chi
Meng, Weibin
Tao, Shimin
Zhao, Xiaofeng
Ma, Hongxia
Zhang, Li
Chen, Boxing
Yang, Hao
Li, Bei
Xiao, Tong
Zhu, Jingbo
Computation and Language
With contributions from the open-source community, a vast amount of instruction tuning (IT) data has emerged. Given the significant resource allocation required for training and evaluating models, it is advantageous to have an efficient method for selecting high-quality IT data. However, existing methods for instruction data selection have limitations such as relying on fragile external APIs, being affected by biases in GPT models, or reducing the diversity of the selected instruction dataset. In this paper, we propose an industrial-friendly, expert-aligned and diversity-preserved instruction data selection method: Clustering and Ranking (CaR). CaR employs a two-step process: first, it ranks instruction pairs using a high-accuracy (84.25%) scoring model aligned with expert preferences; second, it preserves dataset diversity through clustering. In our experiment, CaR efficiently selected a mere 1.96% of Alpaca's IT data, yet the resulting AlpaCaR model surpassed Alpaca's performance by an average of 32.1% in GPT-4 evaluations. Moreover, we find that data selecting is a consistent paradigm whether the pre-trained model is more capable or the model parameters scaling up. Our approach employs compact models with 550M parameters and incurs just 11.2% of the financial outlay of current methods, enhancing its industrial deployability.
title Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality Estimation
topic Computation and Language
url https://arxiv.org/abs/2402.18191