The Lessons of Developing Process Reward Models in Mathematical Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zhenru, Zheng, Chujie, Wu, Yangzhen, Zhang, Beichen, Lin, Runji, Yu, Bowen, Liu, Dayiheng, Zhou, Jingren, Lin, Junyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ProcessBench: Identifying Process Errors in Mathematical Reasoning
von: Zheng, Chujie, et al.
Veröffentlicht: (2024)
von: Zheng, Chujie, et al.
Veröffentlicht: (2024)
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
von: Yang, An, et al.
Veröffentlicht: (2024)
von: Yang, An, et al.
Veröffentlicht: (2024)
START: Self-taught Reasoner with Tools
von: Li, Chengpeng, et al.
Veröffentlicht: (2025)
von: Li, Chengpeng, et al.
Veröffentlicht: (2025)
WorldPM: Scaling Human Preference Modeling
von: Wang, Binghai, et al.
Veröffentlicht: (2025)
von: Wang, Binghai, et al.
Veröffentlicht: (2025)
Chain of Execution Supervision Promotes General Reasoning in Large Language Models
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
Soft Adaptive Policy Optimization
von: Gao, Chang, et al.
Veröffentlicht: (2025)
von: Gao, Chang, et al.
Veröffentlicht: (2025)
Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
von: Wang, Binghai, et al.
Veröffentlicht: (2026)
von: Wang, Binghai, et al.
Veröffentlicht: (2026)
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
von: Zheng, Chujie, et al.
Veröffentlicht: (2025)
von: Zheng, Chujie, et al.
Veröffentlicht: (2025)
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
von: Luo, Ruilin, et al.
Veröffentlicht: (2025)
von: Luo, Ruilin, et al.
Veröffentlicht: (2025)
Group Sequence Policy Optimization
von: Zheng, Chujie, et al.
Veröffentlicht: (2025)
von: Zheng, Chujie, et al.
Veröffentlicht: (2025)
MARGE: Improving Math Reasoning for LLMs with Guided Exploration
von: Gao, Jingyue, et al.
Veröffentlicht: (2025)
von: Gao, Jingyue, et al.
Veröffentlicht: (2025)
Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
von: Wang, Shenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Shenzhi, et al.
Veröffentlicht: (2025)
Language Models can Self-Lengthen to Generate Long Texts
von: Quan, Shanghaoran, et al.
Veröffentlicht: (2024)
von: Quan, Shanghaoran, et al.
Veröffentlicht: (2024)
Disentangling Reasoning Tokens and Boilerplate Tokens For Language Model Fine-tuning
von: Ye, Ziang, et al.
Veröffentlicht: (2024)
von: Ye, Ziang, et al.
Veröffentlicht: (2024)
DataMan: Data Manager for Pre-training Large Language Models
von: Peng, Ru, et al.
Veröffentlicht: (2025)
von: Peng, Ru, et al.
Veröffentlicht: (2025)
An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning
von: Sun, Wei, et al.
Veröffentlicht: (2025)
von: Sun, Wei, et al.
Veröffentlicht: (2025)
Teaching Language Models to Reason with Tools
von: Li, Chengpeng, et al.
Veröffentlicht: (2025)
von: Li, Chengpeng, et al.
Veröffentlicht: (2025)
Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment
von: Lu, Keming, et al.
Veröffentlicht: (2024)
von: Lu, Keming, et al.
Veröffentlicht: (2024)
Revealing Behavioral Plasticity in Large Language Models: A Token-Conditional Perspective
von: Mao, Liyuan, et al.
Veröffentlicht: (2026)
von: Mao, Liyuan, et al.
Veröffentlicht: (2026)
A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning
von: Zhu, Jiachen, et al.
Veröffentlicht: (2025)
von: Zhu, Jiachen, et al.
Veröffentlicht: (2025)
CatCode: A Comprehensive Evaluation Framework for LLMs On the Mixture of Code and Text
von: Lin, Zhenru, et al.
Veröffentlicht: (2024)
von: Lin, Zhenru, et al.
Veröffentlicht: (2024)
Self-Evolving Critique Abilities in Large Language Models
von: Tang, Zhengyang, et al.
Veröffentlicht: (2025)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2025)
RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques
von: Tang, Zhengyang, et al.
Veröffentlicht: (2025)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2025)
JiuZhang3.0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis Models
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
Socratic-PRMBench: Benchmarking Process Reward Models with Systematic Reasoning Patterns
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning
von: Guan, Xin, et al.
Veröffentlicht: (2026)
von: Guan, Xin, et al.
Veröffentlicht: (2026)
RM-R1: Reward Modeling as Reasoning
von: Chen, Xiusi, et al.
Veröffentlicht: (2025)
von: Chen, Xiusi, et al.
Veröffentlicht: (2025)
Evaluating Robustness of Reward Models for Mathematical Reasoning
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
von: Zhang, Yanzhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yanzhao, et al.
Veröffentlicht: (2025)
DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning
von: Li, Chengpeng, et al.
Veröffentlicht: (2024)
von: Li, Chengpeng, et al.
Veröffentlicht: (2024)
Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning
von: Rajaee, Sara, et al.
Veröffentlicht: (2025)
von: Rajaee, Sara, et al.
Veröffentlicht: (2025)
LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback
von: Gao, Bofei, et al.
Veröffentlicht: (2024)
von: Gao, Bofei, et al.
Veröffentlicht: (2024)
ETCHR: Editing To Clarify and Harness Reasoning
von: Zhang, Beichen, et al.
Veröffentlicht: (2026)
von: Zhang, Beichen, et al.
Veröffentlicht: (2026)
ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
von: Fan, Lishui, et al.
Veröffentlicht: (2025)
von: Fan, Lishui, et al.
Veröffentlicht: (2025)
GR-Ben: A General Reasoning Benchmark for Evaluating Process Reward Models
von: Sun, Zhouhao, et al.
Veröffentlicht: (2026)
von: Sun, Zhouhao, et al.
Veröffentlicht: (2026)
CoRT: Code-integrated Reasoning within Thinking
von: Li, Chengpeng, et al.
Veröffentlicht: (2025)
von: Li, Chengpeng, et al.
Veröffentlicht: (2025)
Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
von: Qiu, Zihan, et al.
Veröffentlicht: (2025)
von: Qiu, Zihan, et al.
Veröffentlicht: (2025)
Process-based Self-Rewarding Language Models
von: Zhang, Shimao, et al.
Veröffentlicht: (2025)
von: Zhang, Shimao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ProcessBench: Identifying Process Errors in Mathematical Reasoning
von: Zheng, Chujie, et al.
Veröffentlicht: (2024) -
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
von: Yang, An, et al.
Veröffentlicht: (2024) -
START: Self-taught Reasoner with Tools
von: Li, Chengpeng, et al.
Veröffentlicht: (2025) -
WorldPM: Scaling Human Preference Modeling
von: Wang, Binghai, et al.
Veröffentlicht: (2025) -
Chain of Execution Supervision Promotes General Reasoning in Large Language Models
von: Chen, Nuo, et al.
Veröffentlicht: (2025)