Gradient Imbalance in Direct Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Qinwei, Shi, Jingzhe, Jin, Can, Hwang, Jenq-Neng, Belongie, Serge, Li, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912251812773888
author Ma, Qinwei
Shi, Jingzhe
Jin, Can
Hwang, Jenq-Neng
Belongie, Serge
Li, Lei
author_facet Ma, Qinwei
Shi, Jingzhe
Jin, Can
Hwang, Jenq-Neng
Belongie, Serge
Li, Lei
contents Direct Preference Optimization (DPO) has been proposed as a promising alternative to Proximal Policy Optimization (PPO) based Reinforcement Learning with Human Feedback (RLHF). However, empirical evaluations consistently reveal suboptimal performance in DPO compared to common RLHF pipelines. In this work, we conduct a systematic analysis of DPO's training dynamics and identify gradient imbalance as a critical limitation. We demonstrate theoretically and empirically that this imbalance perturbs optimization trajectories, destabilizes learning, and induces suboptimal convergence. To address this issue, we propose Balanced-DPO, a simple yet effective modification to the DPO objective that introduces a computationally efficient gradient reweighting mechanism. Our experiments demonstrate the effectiveness of Balanced-DPO, validating the theoretical findings and confirming that addressing gradient imbalance is key to improving DPO's performance, highlighting a promising direction for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20847
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gradient Imbalance in Direct Preference Optimization
Ma, Qinwei
Shi, Jingzhe
Jin, Can
Hwang, Jenq-Neng
Belongie, Serge
Li, Lei
Machine Learning
Direct Preference Optimization (DPO) has been proposed as a promising alternative to Proximal Policy Optimization (PPO) based Reinforcement Learning with Human Feedback (RLHF). However, empirical evaluations consistently reveal suboptimal performance in DPO compared to common RLHF pipelines. In this work, we conduct a systematic analysis of DPO's training dynamics and identify gradient imbalance as a critical limitation. We demonstrate theoretically and empirically that this imbalance perturbs optimization trajectories, destabilizes learning, and induces suboptimal convergence. To address this issue, we propose Balanced-DPO, a simple yet effective modification to the DPO objective that introduces a computationally efficient gradient reweighting mechanism. Our experiments demonstrate the effectiveness of Balanced-DPO, validating the theoretical findings and confirming that addressing gradient imbalance is key to improving DPO's performance, highlighting a promising direction for future research.
title Gradient Imbalance in Direct Preference Optimization
topic Machine Learning
url https://arxiv.org/abs/2502.20847