DavIR: Data Selection via Implicit Reward for Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Haotian, Liu, Tingkai, Ma, Qianli, Zhang, Yufeng, Yuan, Jianbo, Liu, Pengfei, You, Yang, Yang, Hongxia
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916531468763136
author Zhou, Haotian
Liu, Tingkai
Ma, Qianli
Zhang, Yufeng
Yuan, Jianbo
Liu, Pengfei
You, Yang
Yang, Hongxia
author_facet Zhou, Haotian
Liu, Tingkai
Ma, Qianli
Zhang, Yufeng
Yuan, Jianbo
Liu, Pengfei
You, Yang
Yang, Hongxia
contents We introduce DavIR, a model-based data selection method for post-training Large Language Models. DavIR generalizes Reducible Holdout Loss to core-set selection problem of causal language modeling, and quantifies the learnability of a given datum with respect to a pre-trained LLM based on relative reduction in loss during fine-tuning, a metric we show to be closely related to the implicit reward model described in Direct Preference Optimization (DPO). We show that 6% of Alpaca dataset selected with DavIR can steer both the LLaMA and Gemma model family to produce superior performance compared to the same models trained on the full 52K dataset. We also show that Alpaca dataset compressed with DavIR can be combined with GSM8K dataset to effectively balance open-domain freeform QA and mathematical reasoning capabilities. Finally, we apply the DavIR objective to DPO and develop a normalized DavIR-DPO objective which improves alignment performance of Zephyr-7B-SFT model by 8% (relative) on AlpacaEval, compared against training on vanilla DPO objective.
format Preprint
id arxiv_https___arxiv_org_abs_2310_13008
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DavIR: Data Selection via Implicit Reward for Large Language Models
Zhou, Haotian
Liu, Tingkai
Ma, Qianli
Zhang, Yufeng
Yuan, Jianbo
Liu, Pengfei
You, Yang
Yang, Hongxia
Machine Learning
Artificial Intelligence
Computation and Language
We introduce DavIR, a model-based data selection method for post-training Large Language Models. DavIR generalizes Reducible Holdout Loss to core-set selection problem of causal language modeling, and quantifies the learnability of a given datum with respect to a pre-trained LLM based on relative reduction in loss during fine-tuning, a metric we show to be closely related to the implicit reward model described in Direct Preference Optimization (DPO). We show that 6% of Alpaca dataset selected with DavIR can steer both the LLaMA and Gemma model family to produce superior performance compared to the same models trained on the full 52K dataset. We also show that Alpaca dataset compressed with DavIR can be combined with GSM8K dataset to effectively balance open-domain freeform QA and mathematical reasoning capabilities. Finally, we apply the DavIR objective to DPO and develop a normalized DavIR-DPO objective which improves alignment performance of Zephyr-7B-SFT model by 8% (relative) on AlpacaEval, compared against training on vanilla DPO objective.
title DavIR: Data Selection via Implicit Reward for Large Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2310.13008