sDPO: Don't Use Your Data All at Once

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Dahyun, Kim, Yungi, Song, Wonho, Kim, Hyeonwoo, Kim, Yunsu, Kim, Sanghoon, Park, Chanjun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909337959530496
author Kim, Dahyun
Kim, Yungi
Song, Wonho
Kim, Hyeonwoo
Kim, Yunsu
Kim, Sanghoon
Park, Chanjun
author_facet Kim, Dahyun
Kim, Yungi
Song, Wonho
Kim, Hyeonwoo
Kim, Yunsu
Kim, Sanghoon
Park, Chanjun
contents As development of large language models (LLM) progresses, aligning them with human preferences has become increasingly important. We propose stepwise DPO (sDPO), an extension of the recently popularized direct preference optimization (DPO) for alignment tuning. This approach involves dividing the available preference datasets and utilizing them in a stepwise manner, rather than employing it all at once. We demonstrate that this method facilitates the use of more precisely aligned reference models within the DPO training framework. Furthermore, sDPO trains the final model to be more performant, even outperforming other popular LLMs with more parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19270
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle sDPO: Don't Use Your Data All at Once
Kim, Dahyun
Kim, Yungi
Song, Wonho
Kim, Hyeonwoo
Kim, Yunsu
Kim, Sanghoon
Park, Chanjun
Computation and Language
Artificial Intelligence
As development of large language models (LLM) progresses, aligning them with human preferences has become increasingly important. We propose stepwise DPO (sDPO), an extension of the recently popularized direct preference optimization (DPO) for alignment tuning. This approach involves dividing the available preference datasets and utilizing them in a stepwise manner, rather than employing it all at once. We demonstrate that this method facilitates the use of more precisely aligned reference models within the DPO training framework. Furthermore, sDPO trains the final model to be more performant, even outperforming other popular LLMs with more parameters.
title sDPO: Don't Use Your Data All at Once
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2403.19270