Linking Process to Outcome: Conditional Reward Modeling for LLM Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zheng, Shan, Ziwei, Song, Kaitao, Li, Yexin, Ren, Kan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Grad2Reward: From Sparse Judgment to Dense Rewards for Improving Open-Ended LLM Reasoning
von: Zhang, Zheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zheng, et al.
Veröffentlicht: (2026)
Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient
von: Li, Changming, et al.
Veröffentlicht: (2026)
von: Li, Changming, et al.
Veröffentlicht: (2026)
Learning to Select In-Context Demonstration Preferred by Large Language Model
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
LLM Reasoning with Process Rewards for Outcome-Guided Steps
von: Rezaei, Mohammad, et al.
Veröffentlicht: (2026)
von: Rezaei, Mohammad, et al.
Veröffentlicht: (2026)
Large Language Models Explore by Latent Distilling
von: Zeng, Yuanhao, et al.
Veröffentlicht: (2026)
von: Zeng, Yuanhao, et al.
Veröffentlicht: (2026)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
von: Zhang, Zhenru, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenru, et al.
Veröffentlicht: (2025)
Prioritize the Process, Not Just the Outcome: Rewarding Latent Thought Trajectories Improves Reasoning in Looped Language Models
von: Williams, Jonathan, et al.
Veröffentlicht: (2026)
von: Williams, Jonathan, et al.
Veröffentlicht: (2026)
Reward Model Generalization for Compute-Aware Test-Time Reasoning
von: Song, Zeen, et al.
Veröffentlicht: (2025)
von: Song, Zeen, et al.
Veröffentlicht: (2025)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
Outcome-based Exploration for LLM Reasoning
von: Song, Yuda, et al.
Veröffentlicht: (2025)
von: Song, Yuda, et al.
Veröffentlicht: (2025)
CAE: Repurposing the Critic as an Explorer in Deep Reinforcement Learning
von: Li, Yexin
Veröffentlicht: (2025)
von: Li, Yexin
Veröffentlicht: (2025)
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
von: Lyu, Chengqi, et al.
Veröffentlicht: (2025)
von: Lyu, Chengqi, et al.
Veröffentlicht: (2025)
Step-wise Rubric Rewards for LLM Reasoning
von: Xie, Weichu, et al.
Veröffentlicht: (2026)
von: Xie, Weichu, et al.
Veröffentlicht: (2026)
EEGFormer: Towards Transferable and Interpretable Large-Scale EEG Foundation Model
von: Chen, Yuqi, et al.
Veröffentlicht: (2024)
von: Chen, Yuqi, et al.
Veröffentlicht: (2024)
Reward-SQL: Boosting Text-to-SQL via Stepwise Reasoning and Process-Supervised Rewards
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025)
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025)
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
von: Cao, Qi, et al.
Veröffentlicht: (2025)
von: Cao, Qi, et al.
Veröffentlicht: (2025)
Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards
von: Fan, Jiajun, et al.
Veröffentlicht: (2025)
von: Fan, Jiajun, et al.
Veröffentlicht: (2025)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
von: Luo, Ruilin, et al.
Veröffentlicht: (2025)
von: Luo, Ruilin, et al.
Veröffentlicht: (2025)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
von: Ding, Ruiyi, et al.
Veröffentlicht: (2026)
von: Ding, Ruiyi, et al.
Veröffentlicht: (2026)
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
von: Ma, Rachel, et al.
Veröffentlicht: (2026)
von: Ma, Rachel, et al.
Veröffentlicht: (2026)
Can Graph Learning Improve Planning in LLM-based Agents?
von: Wu, Xixi, et al.
Veröffentlicht: (2024)
von: Wu, Xixi, et al.
Veröffentlicht: (2024)
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
von: Song, Ruike, et al.
Veröffentlicht: (2025)
von: Song, Ruike, et al.
Veröffentlicht: (2025)
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
von: Zeng, Thomas, et al.
Veröffentlicht: (2025)
von: Zeng, Thomas, et al.
Veröffentlicht: (2025)
FlowRL: Matching Reward Distributions for LLM Reasoning
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation
von: Pan, Pei-Chi, et al.
Veröffentlicht: (2026)
von: Pan, Pei-Chi, et al.
Veröffentlicht: (2026)
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
Process Reward Models for LLM Agents: Practical Framework and Directions
von: Choudhury, Sanjiban
Veröffentlicht: (2025)
von: Choudhury, Sanjiban
Veröffentlicht: (2025)
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
von: Li, Mengqi, et al.
Veröffentlicht: (2025)
von: Li, Mengqi, et al.
Veröffentlicht: (2025)
Unsupervised Process Reward Models
von: Gadetsky, Artyom, et al.
Veröffentlicht: (2026)
von: Gadetsky, Artyom, et al.
Veröffentlicht: (2026)
Accelerating LLM Reasoning via Early Rejection with Partial Reward Modeling
von: Cheshmi, Seyyed Saeid, et al.
Veröffentlicht: (2025)
von: Cheshmi, Seyyed Saeid, et al.
Veröffentlicht: (2025)
Milestones over Outcome: Unlocking Geometric Reasoning with Sub-Goal Verifiable Reward
von: Chen, Jianlong, et al.
Veröffentlicht: (2026)
von: Chen, Jianlong, et al.
Veröffentlicht: (2026)
Rewarding Graph Reasoning Process makes LLMs more Generalized Reasoners
von: Peng, Miao, et al.
Veröffentlicht: (2025)
von: Peng, Miao, et al.
Veröffentlicht: (2025)
Entropy-Regularized Process Reward Model
von: Zhang, Hanning, et al.
Veröffentlicht: (2024)
von: Zhang, Hanning, et al.
Veröffentlicht: (2024)
Conditional Imputation for Within-Modality Missingness in Multi-Modal Federated Learning
von: Zheng, Wugeng, et al.
Veröffentlicht: (2026)
von: Zheng, Wugeng, et al.
Veröffentlicht: (2026)
Smaller Models, Smarter Rewards: A Two-Sided Approach to Process and Outcome Rewards
von: Groeneveld, Jan Niklas, et al.
Veröffentlicht: (2025)
von: Groeneveld, Jan Niklas, et al.
Veröffentlicht: (2025)
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2026)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2026)
Bayesian Reward Models for LLM Alignment
von: Yang, Adam X., et al.
Veröffentlicht: (2024)
von: Yang, Adam X., et al.
Veröffentlicht: (2024)
Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning
von: Cheng, Jie, et al.
Veröffentlicht: (2025)
von: Cheng, Jie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Grad2Reward: From Sparse Judgment to Dense Rewards for Improving Open-Ended LLM Reasoning
von: Zhang, Zheng, et al.
Veröffentlicht: (2026) -
Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient
von: Li, Changming, et al.
Veröffentlicht: (2026) -
Learning to Select In-Context Demonstration Preferred by Large Language Model
von: Zhang, Zheng, et al.
Veröffentlicht: (2025) -
LLM Reasoning with Process Rewards for Outcome-Guided Steps
von: Rezaei, Mohammad, et al.
Veröffentlicht: (2026) -
Large Language Models Explore by Latent Distilling
von: Zeng, Yuanhao, et al.
Veröffentlicht: (2026)