Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Jie, Zhou, Zhanhui, Liu, Jiaheng, Bu, Xingyuan, Yang, Chao, Zhong, Han-Sen, Ouyang, Wanli
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929388254134272
author Liu, Jie
Zhou, Zhanhui
Liu, Jiaheng
Bu, Xingyuan
Yang, Chao
Zhong, Han-Sen
Ouyang, Wanli
author_facet Liu, Jie
Zhou, Zhanhui
Liu, Jiaheng
Bu, Xingyuan
Yang, Chao
Zhong, Han-Sen
Ouyang, Wanli
contents Direct Preference Optimization (DPO), a standard method for aligning language models with human preferences, is traditionally applied to offline preferences. Recent studies show that DPO benefits from iterative training with online preferences labeled by a trained reward model. In this work, we identify a pitfall of vanilla iterative DPO - improved response quality can lead to increased verbosity. To address this, we introduce iterative length-regularized DPO (iLR-DPO) to penalize response length. Our empirical results show that iLR-DPO can enhance a 7B model to perform on par with GPT-4 without increasing verbosity. Specifically, our 7B model achieves a $50.5\%$ length-controlled win rate against $\texttt{GPT-4 Preview}$ on AlpacaEval 2.0, and excels across standard benchmarks including MT-Bench, Arena-Hard and OpenLLM Leaderboard. These results demonstrate the effectiveness of iterative DPO in aligning language models with human feedback.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11817
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level
Liu, Jie
Zhou, Zhanhui
Liu, Jiaheng
Bu, Xingyuan
Yang, Chao
Zhong, Han-Sen
Ouyang, Wanli
Computation and Language
Artificial Intelligence
Machine Learning
Direct Preference Optimization (DPO), a standard method for aligning language models with human preferences, is traditionally applied to offline preferences. Recent studies show that DPO benefits from iterative training with online preferences labeled by a trained reward model. In this work, we identify a pitfall of vanilla iterative DPO - improved response quality can lead to increased verbosity. To address this, we introduce iterative length-regularized DPO (iLR-DPO) to penalize response length. Our empirical results show that iLR-DPO can enhance a 7B model to perform on par with GPT-4 without increasing verbosity. Specifically, our 7B model achieves a $50.5\%$ length-controlled win rate against $\texttt{GPT-4 Preview}$ on AlpacaEval 2.0, and excels across standard benchmarks including MT-Bench, Arena-Hard and OpenLLM Leaderboard. These results demonstrate the effectiveness of iterative DPO in aligning language models with human feedback.
title Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.11817