Gradient Imbalance in Direct Preference Optimization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ma, Qinwei, Shi, Jingzhe, Jin, Can, Hwang, Jenq-Neng, Belongie, Serge, Li, Lei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912251812773888
author Ma, Qinwei
Shi, Jingzhe
Jin, Can
Hwang, Jenq-Neng
Belongie, Serge
Li, Lei
author_facet Ma, Qinwei
Shi, Jingzhe
Jin, Can
Hwang, Jenq-Neng
Belongie, Serge
Li, Lei
contents Direct Preference Optimization (DPO) has been proposed as a promising alternative to Proximal Policy Optimization (PPO) based Reinforcement Learning with Human Feedback (RLHF). However, empirical evaluations consistently reveal suboptimal performance in DPO compared to common RLHF pipelines. In this work, we conduct a systematic analysis of DPO's training dynamics and identify gradient imbalance as a critical limitation. We demonstrate theoretically and empirically that this imbalance perturbs optimization trajectories, destabilizes learning, and induces suboptimal convergence. To address this issue, we propose Balanced-DPO, a simple yet effective modification to the DPO objective that introduces a computationally efficient gradient reweighting mechanism. Our experiments demonstrate the effectiveness of Balanced-DPO, validating the theoretical findings and confirming that addressing gradient imbalance is key to improving DPO's performance, highlighting a promising direction for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20847
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gradient Imbalance in Direct Preference Optimization
Ma, Qinwei
Shi, Jingzhe
Jin, Can
Hwang, Jenq-Neng
Belongie, Serge
Li, Lei
Machine Learning
Direct Preference Optimization (DPO) has been proposed as a promising alternative to Proximal Policy Optimization (PPO) based Reinforcement Learning with Human Feedback (RLHF). However, empirical evaluations consistently reveal suboptimal performance in DPO compared to common RLHF pipelines. In this work, we conduct a systematic analysis of DPO's training dynamics and identify gradient imbalance as a critical limitation. We demonstrate theoretically and empirically that this imbalance perturbs optimization trajectories, destabilizes learning, and induces suboptimal convergence. To address this issue, we propose Balanced-DPO, a simple yet effective modification to the DPO objective that introduces a computationally efficient gradient reweighting mechanism. Our experiments demonstrate the effectiveness of Balanced-DPO, validating the theoretical findings and confirming that addressing gradient imbalance is key to improving DPO's performance, highlighting a promising direction for future research.
title Gradient Imbalance in Direct Preference Optimization
topic Machine Learning
url https://arxiv.org/abs/2502.20847