It Takes Two: Your GRPO Is Secretly DPO
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yihong, Ma, Liheng, Ding, Lei, Li, Muzhi, Wang, Xinyu, Chen, Kejia, Su, Zhan, Zhang, Zhanguang, Huang, Chenyang, Zhang, Yingxue, Coates, Mark, Nie, Jian-Yun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One Demo Is All It Takes: Planning Domain Derivation with LLMs from A Single Demonstration
by: Huang, Jinbang, et al.
Published: (2025)
by: Huang, Jinbang, et al.
Published: (2025)
An Entity Linking Agent for Question Answering
by: Luo, Yajie, et al.
Published: (2025)
by: Luo, Yajie, et al.
Published: (2025)
Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning
by: Wu, Yihong, et al.
Published: (2025)
by: Wu, Yihong, et al.
Published: (2025)
Multi-resolution Time-Series Transformer for Long-term Forecasting
by: Zhang, Yitian, et al.
Published: (2023)
by: Zhang, Yitian, et al.
Published: (2023)
HardCore Generation: Generating Hard UNSAT Problems for Data Augmentation
by: Cotnareanu, Joseph, et al.
Published: (2024)
by: Cotnareanu, Joseph, et al.
Published: (2024)
GraphPPD: Posterior Predictive Modelling for Graph-Level Inference
by: Pal, Soumyasundar, et al.
Published: (2025)
by: Pal, Soumyasundar, et al.
Published: (2025)
Plain Transformers Can be Powerful Graph Learners
by: Ma, Liheng, et al.
Published: (2025)
by: Ma, Liheng, et al.
Published: (2025)
CKGConv: General Graph Convolution with Continuous Kernels
by: Ma, Liheng, et al.
Published: (2024)
by: Ma, Liheng, et al.
Published: (2024)
Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning via Automated Domain Generation
by: Huang, Jinbang, et al.
Published: (2025)
by: Huang, Jinbang, et al.
Published: (2025)
Sparse Decomposition of Graph Neural Networks
by: Hu, Yaochen, et al.
Published: (2024)
by: Hu, Yaochen, et al.
Published: (2024)
From Evidence to Trajectory: Abductive Reasoning Path Synthesis for Training Retrieval-Augmented Generation Agents
by: Li, Muzhi, et al.
Published: (2025)
by: Li, Muzhi, et al.
Published: (2025)
DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory
by: Zhou, Wenxuan, et al.
Published: (2025)
by: Zhou, Wenxuan, et al.
Published: (2025)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
by: Tong, Chengzhuo, et al.
Published: (2025)
by: Tong, Chengzhuo, et al.
Published: (2025)
FEval-TTC: Fair Evaluation Protocol for Test-Time Compute
by: Rumiantsev, Pavel, et al.
Published: (2025)
by: Rumiantsev, Pavel, et al.
Published: (2025)
Refining Answer Distributions for Improved Large Language Model Reasoning
by: Pal, Soumyasundar, et al.
Published: (2024)
by: Pal, Soumyasundar, et al.
Published: (2024)
A Balanced Neuro-Symbolic Approach for Commonsense Abductive Logic
by: Cotnareanu, Joseph, et al.
Published: (2026)
by: Cotnareanu, Joseph, et al.
Published: (2026)
Personalized Negative Reservoir for Incremental Learning in Recommender Systems
by: Valkanas, Antonios, et al.
Published: (2024)
by: Valkanas, Antonios, et al.
Published: (2024)
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
by: Dai, Muzhi, et al.
Published: (2025)
by: Dai, Muzhi, et al.
Published: (2025)
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting
by: Zhang, Yitian, et al.
Published: (2025)
by: Zhang, Yitian, et al.
Published: (2025)
GraSS: Combining Graph Neural Networks with Expert Knowledge for SAT Solver Selection
by: Zhang, Zhanguang, et al.
Published: (2024)
by: Zhang, Zhanguang, et al.
Published: (2024)
Enhancing Logical Reasoning in Large Language Models through Graph-based Synthetic Data
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
DyG2Vec: Efficient Representation Learning for Dynamic Graphs
by: Alomrani, Mohammad Ali, et al.
Published: (2022)
by: Alomrani, Mohammad Ali, et al.
Published: (2022)
Evaluating GRPO and DPO for Faithful Chain-of-Thought Reasoning in LLMs
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
E2ESlack: An End-to-End Graph-Based Framework for Pre-Routing Slack Prediction
by: Bodhe, Saurabh, et al.
Published: (2025)
by: Bodhe, Saurabh, et al.
Published: (2025)
C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for Reasoning
by: Valkanas, Antonios, et al.
Published: (2025)
by: Valkanas, Antonios, et al.
Published: (2025)
InnerThoughts: Disentangling Representations and Predictions in Large Language Models
by: Chételat, Didier, et al.
Published: (2025)
by: Chételat, Didier, et al.
Published: (2025)
H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model
by: Huang, Jinbang, et al.
Published: (2026)
by: Huang, Jinbang, et al.
Published: (2026)
CodeDPO: Aligning Code Models with Self Generated and Verified Source Code
by: Zhang, Kechi, et al.
Published: (2024)
by: Zhang, Kechi, et al.
Published: (2024)
Exploring the Best Practices of Query Expansion with Large Language Models
by: Zhang, Le, et al.
Published: (2024)
by: Zhang, Le, et al.
Published: (2024)
GRPO is Secretly a Process Reward Model
by: Sullivan, Michael, et al.
Published: (2025)
by: Sullivan, Michael, et al.
Published: (2025)
GIFT: Group-Relative Implicit Fine-Tuning Integrates GRPO with DPO and UNA
by: Wang, Zhichao
Published: (2025)
by: Wang, Zhichao
Published: (2025)
Omni-Thinker: Scaling Multi-Task RL in LLMs with Hybrid Reward and Task Scheduling
by: Li, Derek, et al.
Published: (2025)
by: Li, Derek, et al.
Published: (2025)
Abductive Reasoning with Probabilistic Commonsense
by: Cotnareanu, Joseph, et al.
Published: (2026)
by: Cotnareanu, Joseph, et al.
Published: (2026)
Aligning Query Representation with Rewritten Query and Relevance Judgments in Conversational Search
by: Mo, Fengran, et al.
Published: (2024)
by: Mo, Fengran, et al.
Published: (2024)
Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
by: Yao, Chaorui, et al.
Published: (2025)
by: Yao, Chaorui, et al.
Published: (2025)
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Unifying Graph Convolution and Contrastive Learning in Collaborative Filtering
by: Wu, Yihong, et al.
Published: (2024)
by: Wu, Yihong, et al.
Published: (2024)
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
by: Chen, Xiwen, et al.
Published: (2025)
by: Chen, Xiwen, et al.
Published: (2025)
Similar Items
-
One Demo Is All It Takes: Planning Domain Derivation with LLMs from A Single Demonstration
by: Huang, Jinbang, et al.
Published: (2025) -
An Entity Linking Agent for Question Answering
by: Luo, Yajie, et al.
Published: (2025) -
Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning
by: Wu, Yihong, et al.
Published: (2025) -
Multi-resolution Time-Series Transformer for Long-term Forecasting
by: Zhang, Yitian, et al.
Published: (2023) -
HardCore Generation: Generating Hard UNSAT Problems for Data Augmentation
by: Cotnareanu, Joseph, et al.
Published: (2024)