DPO Meets PPO: Reinforced Token Optimization for RLHF
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhong, Han, Shan, Zikang, Feng, Guhao, Xiong, Wei, Cheng, Xinle, Zhao, Li, He, Di, Bian, Jiang, Wang, Liwei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning
di: Shan, Zikang, et al.
Pubblicazione: (2026)
di: Shan, Zikang, et al.
Pubblicazione: (2026)
Theoretical Benefit and Limitation of Diffusion Language Model
di: Feng, Guhao, et al.
Pubblicazione: (2025)
di: Feng, Guhao, et al.
Pubblicazione: (2025)
Rethinking Model-based, Policy-based, and Value-based Reinforcement Learning via the Lens of Representation Complexity
di: Feng, Guhao, et al.
Pubblicazione: (2023)
di: Feng, Guhao, et al.
Pubblicazione: (2023)
RLHF Workflow: From Reward Modeling to Online RLHF
di: Dong, Hanze, et al.
Pubblicazione: (2024)
di: Dong, Hanze, et al.
Pubblicazione: (2024)
Do Efficient Transformers Really Save Computation?
di: Yang, Kai, et al.
Pubblicazione: (2024)
di: Yang, Kai, et al.
Pubblicazione: (2024)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
di: He, Zhenyu, et al.
Pubblicazione: (2024)
di: He, Zhenyu, et al.
Pubblicazione: (2024)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
di: He, Chaoyue, et al.
Pubblicazione: (2026)
di: He, Chaoyue, et al.
Pubblicazione: (2026)
How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs
di: Feng, Guhao, et al.
Pubblicazione: (2024)
di: Feng, Guhao, et al.
Pubblicazione: (2024)
Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression
di: Huang, Jiameng, et al.
Pubblicazione: (2025)
di: Huang, Jiameng, et al.
Pubblicazione: (2025)
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
di: Xu, Shusheng, et al.
Pubblicazione: (2024)
di: Xu, Shusheng, et al.
Pubblicazione: (2024)
In-Place Test-Time Training
di: Feng, Guhao, et al.
Pubblicazione: (2026)
di: Feng, Guhao, et al.
Pubblicazione: (2026)
Reward Difference Optimization For Sample Reweighting In Offline RLHF
di: Wang, Shiqi, et al.
Pubblicazione: (2024)
di: Wang, Shiqi, et al.
Pubblicazione: (2024)
APRIL: API Synthesis with Automatic Prompt Optimization and Reinforcement Learning
di: Zhong, Hua, et al.
Pubblicazione: (2025)
di: Zhong, Hua, et al.
Pubblicazione: (2025)
WPO: Enhancing RLHF with Weighted Preference Optimization
di: Zhou, Wenxuan, et al.
Pubblicazione: (2024)
di: Zhou, Wenxuan, et al.
Pubblicazione: (2024)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
di: Wu, Junkang, et al.
Pubblicazione: (2024)
di: Wu, Junkang, et al.
Pubblicazione: (2024)
Dataset Reset Policy Optimization for RLHF
di: Chang, Jonathan D., et al.
Pubblicazione: (2024)
di: Chang, Jonathan D., et al.
Pubblicazione: (2024)
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
di: Li, Shilong, et al.
Pubblicazione: (2024)
di: Li, Shilong, et al.
Pubblicazione: (2024)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
di: Hu, Jian, et al.
Pubblicazione: (2024)
di: Hu, Jian, et al.
Pubblicazione: (2024)
VidTok: A Versatile and Open-Source Video Tokenizer
di: Tang, Anni, et al.
Pubblicazione: (2024)
di: Tang, Anni, et al.
Pubblicazione: (2024)
Reinforcement Learning with Token-level Feedback for Controllable Text Generation
di: Li, Wendi, et al.
Pubblicazione: (2024)
di: Li, Wendi, et al.
Pubblicazione: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
The Perfect Blend: Redefining RLHF with Mixture of Judges
di: Xu, Tengyu, et al.
Pubblicazione: (2024)
di: Xu, Tengyu, et al.
Pubblicazione: (2024)
BinaryPPO: Efficient Policy Optimization for Binary Classification
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)
Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment
di: Yang, Zhiqin, et al.
Pubblicazione: (2026)
di: Yang, Zhiqin, et al.
Pubblicazione: (2026)
DPO-Shift: Shifting the Distribution of Direct Preference Optimization
di: Yang, Xiliang, et al.
Pubblicazione: (2025)
di: Yang, Xiliang, et al.
Pubblicazione: (2025)
Active Preference Optimization for Sample Efficient RLHF
di: Das, Nirjhar, et al.
Pubblicazione: (2024)
di: Das, Nirjhar, et al.
Pubblicazione: (2024)
An Empirical Study of SFT-DPO Interaction and Parameterization in Small Language Models
di: Feng, Yuming, et al.
Pubblicazione: (2026)
di: Feng, Yuming, et al.
Pubblicazione: (2026)
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
di: Ivison, Hamish, et al.
Pubblicazione: (2024)
di: Ivison, Hamish, et al.
Pubblicazione: (2024)
Cat-DPO: Category-Adaptive Safety Alignment
di: Yang, Tiankai, et al.
Pubblicazione: (2026)
di: Yang, Tiankai, et al.
Pubblicazione: (2026)
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
di: Zhong, Yiwu, et al.
Pubblicazione: (2024)
di: Zhong, Yiwu, et al.
Pubblicazione: (2024)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
di: Zhang, Zhengze, et al.
Pubblicazione: (2025)
di: Zhang, Zhengze, et al.
Pubblicazione: (2025)
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
di: Feng, Duanyu, et al.
Pubblicazione: (2024)
di: Feng, Duanyu, et al.
Pubblicazione: (2024)
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
di: Shi, Ruizhe, et al.
Pubblicazione: (2025)
di: Shi, Ruizhe, et al.
Pubblicazione: (2025)
MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
di: Lan, Guangchen, et al.
Pubblicazione: (2025)
di: Lan, Guangchen, et al.
Pubblicazione: (2025)
Playing with Transformer at 30+ FPS via Next-Frame Diffusion
di: Cheng, Xinle, et al.
Pubblicazione: (2025)
di: Cheng, Xinle, et al.
Pubblicazione: (2025)
Distributionally Robust Token Optimization in RLHF
di: Jin, Yeping, et al.
Pubblicazione: (2026)
di: Jin, Yeping, et al.
Pubblicazione: (2026)
Context-DPO: Aligning Language Models for Context-Faithfulness
di: Bi, Baolong, et al.
Pubblicazione: (2024)
di: Bi, Baolong, et al.
Pubblicazione: (2024)
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
di: Qi, Biqing, et al.
Pubblicazione: (2024)
di: Qi, Biqing, et al.
Pubblicazione: (2024)
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
di: Mohamed, Anas, et al.
Pubblicazione: (2025)
di: Mohamed, Anas, et al.
Pubblicazione: (2025)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
di: Gupta, Taneesh, et al.
Pubblicazione: (2024)
di: Gupta, Taneesh, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning
di: Shan, Zikang, et al.
Pubblicazione: (2026) -
Theoretical Benefit and Limitation of Diffusion Language Model
di: Feng, Guhao, et al.
Pubblicazione: (2025) -
Rethinking Model-based, Policy-based, and Value-based Reinforcement Learning via the Lens of Representation Complexity
di: Feng, Guhao, et al.
Pubblicazione: (2023) -
RLHF Workflow: From Reward Modeling to Online RLHF
di: Dong, Hanze, et al.
Pubblicazione: (2024) -
Do Efficient Transformers Really Save Computation?
di: Yang, Kai, et al.
Pubblicazione: (2024)