Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Xuerui, Wang, Yue, Zhu, Jinhua, Yi, Mingyang, Xu, Feng, Ma, Zhiming, Liu, Yuting |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
by: He, Chaoyue, et al.
Published: (2026)
by: He, Chaoyue, et al.
Published: (2026)
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
by: Zhang, Zhengze, et al.
Published: (2025)
by: Zhang, Zhengze, et al.
Published: (2025)
It Takes Two: Your GRPO Is Secretly DPO
by: Wu, Yihong, et al.
Published: (2025)
by: Wu, Yihong, et al.
Published: (2025)
What Matters in Data for DPO?
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
by: Wang, Bo, et al.
Published: (2025)
by: Wang, Bo, et al.
Published: (2025)
Bootstrapping Language Models with DPO Implicit Rewards
by: Chen, Changyu, et al.
Published: (2024)
by: Chen, Changyu, et al.
Published: (2024)
Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
by: Wen, Liang, et al.
Published: (2025)
by: Wen, Liang, et al.
Published: (2025)
Preference Robustness for DPO with Applications to Public Health
by: Kim, Cheol Woo, et al.
Published: (2025)
by: Kim, Cheol Woo, et al.
Published: (2025)
DPO Meets PPO: Reinforced Token Optimization for RLHF
by: Zhong, Han, et al.
Published: (2024)
by: Zhong, Han, et al.
Published: (2024)
Why DPO is a Misspecified Estimator and How to Fix It
by: Gopalan, Aditya, et al.
Published: (2025)
by: Gopalan, Aditya, et al.
Published: (2025)
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
by: Pipano, Idan, et al.
Published: (2026)
by: Pipano, Idan, et al.
Published: (2026)
CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment
by: Liu, Jilong, et al.
Published: (2026)
by: Liu, Jilong, et al.
Published: (2026)
Hard Negative Sample-Augmented DPO Post-Training for Small Language Models
by: Lu, Haocheng, et al.
Published: (2025)
by: Lu, Haocheng, et al.
Published: (2025)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
by: Tong, Chengzhuo, et al.
Published: (2025)
by: Tong, Chengzhuo, et al.
Published: (2025)
$β$-DPO: Direct Preference Optimization with Dynamic $β$
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
C2-DPO: Constrained Controlled Direct Preference Optimization
by: Asadi, Kavosh, et al.
Published: (2025)
by: Asadi, Kavosh, et al.
Published: (2025)
Provably Convergent Primal-Dual DPO for Constrained LLM Alignment
by: Du, Yihan, et al.
Published: (2025)
by: Du, Yihan, et al.
Published: (2025)
g-DPO: Scalable Preference Optimization for Protein Language Models
by: Ferragu, Constance, et al.
Published: (2025)
by: Ferragu, Constance, et al.
Published: (2025)
Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models
by: Mouiche, Inoussa
Published: (2026)
by: Mouiche, Inoussa
Published: (2026)
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
by: Xiao, Teng, et al.
Published: (2024)
by: Xiao, Teng, et al.
Published: (2024)
Robust Multi-Objective Preference Alignment with Online DPO
by: Gupta, Raghav, et al.
Published: (2025)
by: Gupta, Raghav, et al.
Published: (2025)
The Viscosity of Logic: Phase Transitions and Hysteresis in DPO Alignment
by: Pollanen, Marco
Published: (2026)
by: Pollanen, Marco
Published: (2026)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
by: Qi, Xuan, et al.
Published: (2025)
by: Qi, Xuan, et al.
Published: (2025)
InCo-DPO: Balancing Distribution Shift and Data Quality for Enhanced Preference Optimization
by: Wang, Yunan, et al.
Published: (2025)
by: Wang, Yunan, et al.
Published: (2025)
MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization
by: Surana, Rohan, et al.
Published: (2026)
by: Surana, Rohan, et al.
Published: (2026)
GIFT: Group-Relative Implicit Fine-Tuning Integrates GRPO with DPO and UNA
by: Wang, Zhichao
Published: (2025)
by: Wang, Zhichao
Published: (2025)
Provably Robust DPO: Aligning Language Models with Noisy Feedback
by: Chowdhury, Sayak Ray, et al.
Published: (2024)
by: Chowdhury, Sayak Ray, et al.
Published: (2024)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
by: Imai, Saki, et al.
Published: (2026)
by: Imai, Saki, et al.
Published: (2026)
Aligning Compound AI Systems via System-level DPO
by: Wang, Xiangwen, et al.
Published: (2025)
by: Wang, Xiangwen, et al.
Published: (2025)
DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization
by: Zhou, Zhenglin, et al.
Published: (2025)
by: Zhou, Zhenglin, et al.
Published: (2025)
Can DPO Learn Diverse Human Values? A Theoretical Scaling Law
by: Im, Shawn, et al.
Published: (2024)
by: Im, Shawn, et al.
Published: (2024)
Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
by: Shi, Ruizhe, et al.
Published: (2025)
by: Shi, Ruizhe, et al.
Published: (2025)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
$ξ$-DPO: Direct Preference Optimization via Ratio Reward Margin
by: Fan, Zhengyuan, et al.
Published: (2026)
by: Fan, Zhengyuan, et al.
Published: (2026)
Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs
by: Oh, Giyeong, et al.
Published: (2026)
by: Oh, Giyeong, et al.
Published: (2026)
SEE-DPO: Self Entropy Enhanced Direct Preference Optimization
by: Shekhar, Shivanshu, et al.
Published: (2024)
by: Shekhar, Shivanshu, et al.
Published: (2024)
Similar Items
-
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
by: He, Chaoyue, et al.
Published: (2026) -
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
by: Su, Xuerui, et al.
Published: (2025) -
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
by: Zhang, Zhengze, et al.
Published: (2025) -
It Takes Two: Your GRPO Is Secretly DPO
by: Wu, Yihong, et al.
Published: (2025) -
What Matters in Data for DPO?
by: Pan, Yu, et al.
Published: (2025)