Saved in:
Bibliographic Details
Main Authors: Liu, Shunyu, Fang, Wenkai, Hu, Zetian, Zhang, Junjie, Zhou, Yang, Zhang, Kongcheng, Tu, Rongcheng, Lin, Ting-En, Huang, Fei, Song, Mingli, Li, Yongbin, Tao, Dacheng
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2503.11701
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913737823223808
author Liu, Shunyu
Fang, Wenkai
Hu, Zetian
Zhang, Junjie
Zhou, Yang
Zhang, Kongcheng
Tu, Rongcheng
Lin, Ting-En
Huang, Fei
Song, Mingli
Li, Yongbin
Tao, Dacheng
author_facet Liu, Shunyu
Fang, Wenkai
Hu, Zetian
Zhang, Junjie
Zhou, Yang
Zhang, Kongcheng
Tu, Rongcheng
Lin, Ting-En
Huang, Fei
Song, Mingli
Li, Yongbin
Tao, Dacheng
contents Large Language Models (LLMs) have demonstrated unprecedented generative capabilities, yet their alignment with human values remains critical for ensuring helpful and harmless deployments. While Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful paradigm for aligning LLMs with human preferences, its reliance on complex reward modeling introduces inherent trade-offs in computational efficiency and training stability. In this context, Direct Preference Optimization (DPO) has recently gained prominence as a streamlined alternative that directly optimizes LLMs using human preferences, thereby circumventing the need for explicit reward modeling. Owing to its theoretical elegance and computational efficiency, DPO has rapidly attracted substantial research efforts exploring its various implementations and applications. However, this field currently lacks systematic organization and comparative analysis. In this survey, we conduct a comprehensive overview of DPO and introduce a novel taxonomy, categorizing previous works into four key dimensions: data strategy, learning framework, constraint mechanism, and model property. We further present a rigorous empirical analysis of DPO variants across standardized benchmarks. Additionally, we discuss real-world applications, open challenges, and future directions for DPO. This work delivers both a conceptual framework for understanding DPO and practical guidance for practitioners, aiming to advance robust and generalizable alignment paradigms. All collected resources are available and will be continuously updated at https://github.com/liushunyu/awesome-direct-preference-optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11701
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey of Direct Preference Optimization
Liu, Shunyu
Fang, Wenkai
Hu, Zetian
Zhang, Junjie
Zhou, Yang
Zhang, Kongcheng
Tu, Rongcheng
Lin, Ting-En
Huang, Fei
Song, Mingli
Li, Yongbin
Tao, Dacheng
Machine Learning
Large Language Models (LLMs) have demonstrated unprecedented generative capabilities, yet their alignment with human values remains critical for ensuring helpful and harmless deployments. While Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful paradigm for aligning LLMs with human preferences, its reliance on complex reward modeling introduces inherent trade-offs in computational efficiency and training stability. In this context, Direct Preference Optimization (DPO) has recently gained prominence as a streamlined alternative that directly optimizes LLMs using human preferences, thereby circumventing the need for explicit reward modeling. Owing to its theoretical elegance and computational efficiency, DPO has rapidly attracted substantial research efforts exploring its various implementations and applications. However, this field currently lacks systematic organization and comparative analysis. In this survey, we conduct a comprehensive overview of DPO and introduce a novel taxonomy, categorizing previous works into four key dimensions: data strategy, learning framework, constraint mechanism, and model property. We further present a rigorous empirical analysis of DPO variants across standardized benchmarks. Additionally, we discuss real-world applications, open challenges, and future directions for DPO. This work delivers both a conceptual framework for understanding DPO and practical guidance for practitioners, aiming to advance robust and generalizable alignment paradigms. All collected resources are available and will be continuously updated at https://github.com/liushunyu/awesome-direct-preference-optimization.
title A Survey of Direct Preference Optimization
topic Machine Learning
url https://arxiv.org/abs/2503.11701