Improving Data Efficiency via Curating LLM-Driven Rating Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pang, Jinlong, Wei, Jiaheng, Shah, Ankit Parag, Zhu, Zhaowei, Wang, Yaxuan, Qian, Chen, Liu, Yang, Bao, Yujia, Wei, Wei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929743789555712
author Pang, Jinlong
Wei, Jiaheng
Shah, Ankit Parag
Zhu, Zhaowei
Wang, Yaxuan
Qian, Chen
Liu, Yang
Bao, Yujia
Wei, Wei
author_facet Pang, Jinlong
Wei, Jiaheng
Shah, Ankit Parag
Zhu, Zhaowei
Wang, Yaxuan
Qian, Chen
Liu, Yang
Bao, Yujia
Wei, Wei
contents Instruction tuning is critical for adapting large language models (LLMs) to downstream tasks, and recent studies have demonstrated that small amounts of human-curated data can outperform larger datasets, challenging traditional data scaling laws. While LLM-based data quality rating systems offer a cost-effective alternative to human annotation, they often suffer from inaccuracies and biases, even in powerful models like GPT-4. In this work, we introduce DS2, a Diversity-aware Score curation method for Data Selection. By systematically modeling error patterns through a score transition matrix, DS2 corrects LLM-based scores and promotes diversity in the selected data samples. Our approach shows that a curated subset (just 3.3% of the original dataset) outperforms full-scale datasets (300k samples) across various machine-alignment benchmarks, and matches or surpasses human-aligned datasets such as LIMA with the same sample size (1k samples). These findings challenge conventional data scaling assumptions, highlighting that redundant, low-quality samples can degrade performance and reaffirming that "more can be less."
format Preprint
id arxiv_https___arxiv_org_abs_2410_10877
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Data Efficiency via Curating LLM-Driven Rating Systems
Pang, Jinlong
Wei, Jiaheng
Shah, Ankit Parag
Zhu, Zhaowei
Wang, Yaxuan
Qian, Chen
Liu, Yang
Bao, Yujia
Wei, Wei
Computation and Language
Artificial Intelligence
Instruction tuning is critical for adapting large language models (LLMs) to downstream tasks, and recent studies have demonstrated that small amounts of human-curated data can outperform larger datasets, challenging traditional data scaling laws. While LLM-based data quality rating systems offer a cost-effective alternative to human annotation, they often suffer from inaccuracies and biases, even in powerful models like GPT-4. In this work, we introduce DS2, a Diversity-aware Score curation method for Data Selection. By systematically modeling error patterns through a score transition matrix, DS2 corrects LLM-based scores and promotes diversity in the selected data samples. Our approach shows that a curated subset (just 3.3% of the original dataset) outperforms full-scale datasets (300k samples) across various machine-alignment benchmarks, and matches or surpasses human-aligned datasets such as LIMA with the same sample size (1k samples). These findings challenge conventional data scaling assumptions, highlighting that redundant, low-quality samples can degrade performance and reaffirming that "more can be less."
title Improving Data Efficiency via Curating LLM-Driven Rating Systems
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.10877