Pre-DPO: Improving Data Utilization in Direct Preference Optimization Using a Guiding Reference Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Junshu, Shen, Wei, Huang, Shulin, Zhou, Qiji, Zhang, Yue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908739102048256
author Pan, Junshu
Shen, Wei
Huang, Shulin
Zhou, Qiji
Zhang, Yue
author_facet Pan, Junshu
Shen, Wei
Huang, Shulin
Zhou, Qiji
Zhang, Yue
contents Direct Preference Optimization (DPO) simplifies reinforcement learning from human feedback (RLHF) for large language models (LLMs) by directly optimizing human preferences without an explicit reward model. We find that during DPO training, the reference model plays the role of a data weight adjuster. However, the common practice of initializing the policy and reference models identically in DPO can lead to inefficient data utilization and impose a performance ceiling. Meanwhile, the lack of a reference model in Simple Preference Optimization (SimPO) reduces training robustness and necessitates stricter conditions to prevent catastrophic forgetting. In this work, we propose Pre-DPO, a simple yet effective DPO-based training paradigm that enhances preference optimization performance by leveraging a guiding reference model. This reference model provides foresight into the optimal policy state achievable through the training preference data, serving as a guiding mechanism that adaptively assigns higher weights to samples more suitable for the model and lower weights to those less suitable. Extensive experiments on AlpacaEval 2.0 and Arena-Hard v0.1 benchmarks demonstrate that Pre-DPO consistently improves the performance of both DPO and SimPO, without relying on external models or additional data.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15843
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pre-DPO: Improving Data Utilization in Direct Preference Optimization Using a Guiding Reference Model
Pan, Junshu
Shen, Wei
Huang, Shulin
Zhou, Qiji
Zhang, Yue
Computation and Language
Direct Preference Optimization (DPO) simplifies reinforcement learning from human feedback (RLHF) for large language models (LLMs) by directly optimizing human preferences without an explicit reward model. We find that during DPO training, the reference model plays the role of a data weight adjuster. However, the common practice of initializing the policy and reference models identically in DPO can lead to inefficient data utilization and impose a performance ceiling. Meanwhile, the lack of a reference model in Simple Preference Optimization (SimPO) reduces training robustness and necessitates stricter conditions to prevent catastrophic forgetting. In this work, we propose Pre-DPO, a simple yet effective DPO-based training paradigm that enhances preference optimization performance by leveraging a guiding reference model. This reference model provides foresight into the optimal policy state achievable through the training preference data, serving as a guiding mechanism that adaptively assigns higher weights to samples more suitable for the model and lower weights to those less suitable. Extensive experiments on AlpacaEval 2.0 and Arena-Hard v0.1 benchmarks demonstrate that Pre-DPO consistently improves the performance of both DPO and SimPO, without relying on external models or additional data.
title Pre-DPO: Improving Data Utilization in Direct Preference Optimization Using a Guiding Reference Model
topic Computation and Language
url https://arxiv.org/abs/2504.15843