Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Xuerui, Wang, Yue, Zhu, Jinhua, Yi, Mingyang, Xu, Feng, Ma, Zhiming, Liu, Yuting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910815146213376
author Su, Xuerui
Wang, Yue
Zhu, Jinhua
Yi, Mingyang
Xu, Feng
Ma, Zhiming
Liu, Yuting
author_facet Su, Xuerui
Wang, Yue
Zhu, Jinhua
Yi, Mingyang
Xu, Feng
Ma, Zhiming
Liu, Yuting
contents With the rapid development of Large Language Models (LLMs), numerous Reinforcement Learning from Human Feedback (RLHF) algorithms have been introduced to improve model safety and alignment with human preferences. These algorithms can be divided into two main frameworks based on whether they require an explicit reward (or value) function for training: actor-critic-based Proximal Policy Optimization (PPO) and alignment-based Direct Preference Optimization (DPO). The mismatch between DPO and PPO, such as DPO's use of a classification loss driven by human-preferred data, has raised confusion about whether DPO should be classified as a Reinforcement Learning (RL) algorithm. To address these ambiguities, we focus on three key aspects related to DPO, RL, and other RLHF algorithms: (1) the construction of the loss function; (2) the target distribution at which the algorithm converges; (3) the impact of key components within the loss function. Specifically, we first establish a unified framework named UDRRA connecting these algorithms based on the construction of their loss functions. Next, we uncover their target policy distributions within this framework. Finally, we investigate the critical components of DPO to understand their impact on the convergence rate. Our work provides a deeper understanding of the relationship between DPO, RL, and other RLHF algorithms, offering new insights for improving existing algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2502_03095
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
Su, Xuerui
Wang, Yue
Zhu, Jinhua
Yi, Mingyang
Xu, Feng
Ma, Zhiming
Liu, Yuting
Machine Learning
With the rapid development of Large Language Models (LLMs), numerous Reinforcement Learning from Human Feedback (RLHF) algorithms have been introduced to improve model safety and alignment with human preferences. These algorithms can be divided into two main frameworks based on whether they require an explicit reward (or value) function for training: actor-critic-based Proximal Policy Optimization (PPO) and alignment-based Direct Preference Optimization (DPO). The mismatch between DPO and PPO, such as DPO's use of a classification loss driven by human-preferred data, has raised confusion about whether DPO should be classified as a Reinforcement Learning (RL) algorithm. To address these ambiguities, we focus on three key aspects related to DPO, RL, and other RLHF algorithms: (1) the construction of the loss function; (2) the target distribution at which the algorithm converges; (3) the impact of key components within the loss function. Specifically, we first establish a unified framework named UDRRA connecting these algorithms based on the construction of their loss functions. Next, we uncover their target policy distributions within this framework. Finally, we investigate the critical components of DPO to understand their impact on the convergence rate. Our work provides a deeper understanding of the relationship between DPO, RL, and other RLHF algorithms, offering new insights for improving existing algorithms.
title Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
topic Machine Learning
url https://arxiv.org/abs/2502.03095