Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Yulan, Ouyang, Sheng, Zhao, Jinman, Liu, Yong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
by: Ouyang, Sheng, et al.
Published: (2025)
by: Ouyang, Sheng, et al.
Published: (2025)
Towards Comprehensive Preference Data Collection for Reward Modeling
by: Hu, Yulan, et al.
Published: (2024)
by: Hu, Yulan, et al.
Published: (2024)
GUNDAM: Aligning Large Language Models with Graph Understanding
by: Ouyang, Sheng, et al.
Published: (2024)
by: Ouyang, Sheng, et al.
Published: (2024)
Refining Latent Representations: A Generative SSL Approach for Heterogeneous Graph Learning
by: Hu, Yulan, et al.
Published: (2023)
by: Hu, Yulan, et al.
Published: (2023)
VIGraph: Generative Self-supervised Learning for Class-Imbalanced Node Classification
by: Hu, Yulan, et al.
Published: (2023)
by: Hu, Yulan, et al.
Published: (2023)
Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning
by: Zhu, Jiachen, et al.
Published: (2025)
by: Zhu, Jiachen, et al.
Published: (2025)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
by: Zhang, Zhenru, et al.
Published: (2025)
by: Zhang, Zhenru, et al.
Published: (2025)
Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning
by: Yang, Zhaohui, et al.
Published: (2025)
by: Yang, Zhaohui, et al.
Published: (2025)
Preserving Node Distinctness in Graph Autoencoders via Similarity Distillation
by: Chen, Ge, et al.
Published: (2024)
by: Chen, Ge, et al.
Published: (2024)
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
by: Zheng, Congmin, et al.
Published: (2025)
by: Zheng, Congmin, et al.
Published: (2025)
Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR
by: Yi, Hao, et al.
Published: (2026)
by: Yi, Hao, et al.
Published: (2026)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
by: Luo, Ruilin, et al.
Published: (2025)
by: Luo, Ruilin, et al.
Published: (2025)
GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models
by: Wang, Zhijie
Published: (2026)
by: Wang, Zhijie
Published: (2026)
What Are Step-Level Reward Models Rewarding? Counterintuitive Findings from MCTS-Boosted Mathematical Reasoning
by: Ma, Yiran, et al.
Published: (2024)
by: Ma, Yiran, et al.
Published: (2024)
An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning
by: Sun, Wei, et al.
Published: (2025)
by: Sun, Wei, et al.
Published: (2025)
Evaluating Robustness of Reward Models for Mathematical Reasoning
by: Kim, Sunghwan, et al.
Published: (2024)
by: Kim, Sunghwan, et al.
Published: (2024)
Exploring Task Unification in Graph Representation Learning via Generative Approach
by: Hu, Yulan, et al.
Published: (2024)
by: Hu, Yulan, et al.
Published: (2024)
Efficient Reasoning via Reward Model
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning
by: Rajaee, Sara, et al.
Published: (2025)
by: Rajaee, Sara, et al.
Published: (2025)
Socratic-PRMBench: Benchmarking Process Reward Models with Systematic Reasoning Patterns
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
Verifiable Process Rewards for Agentic Reasoning
by: Yuan, Huining, et al.
Published: (2026)
by: Yuan, Huining, et al.
Published: (2026)
SPPD: Self-training with Process Preference Learning Using Dynamic Value Margin
by: Yi, Hao, et al.
Published: (2025)
by: Yi, Hao, et al.
Published: (2025)
Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical Reasoning
by: Han, Jiuzhou, et al.
Published: (2025)
by: Han, Jiuzhou, et al.
Published: (2025)
Advancing Reasoning in Diffusion Language Models with Denoising Process Rewards
by: Xie, Shaoan, et al.
Published: (2025)
by: Xie, Shaoan, et al.
Published: (2025)
Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling
by: Wang, Jiaxuan, et al.
Published: (2026)
by: Wang, Jiaxuan, et al.
Published: (2026)
LogReasoner: Empowering LLMs with Expert-like Coarse-to-Fine Reasoning for Automated Log Analysis
by: Ma, Lipeng, et al.
Published: (2025)
by: Ma, Lipeng, et al.
Published: (2025)
Dirichlet-Based Coarse-to-Fine Example Selection For Open-Set Annotation
by: Wang, Ye-Wen, et al.
Published: (2024)
by: Wang, Ye-Wen, et al.
Published: (2024)
MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
by: Chen, Jinhao, et al.
Published: (2025)
by: Chen, Jinhao, et al.
Published: (2025)
GR-Ben: A General Reasoning Benchmark for Evaluating Process Reward Models
by: Sun, Zhouhao, et al.
Published: (2026)
by: Sun, Zhouhao, et al.
Published: (2026)
Distilling Mathematical Reasoning Capabilities into Small Language Models
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks
by: Cao, Linbo, et al.
Published: (2025)
by: Cao, Linbo, et al.
Published: (2025)
RPM-MCTS: Knowledge-Retrieval as Process Reward Model with Monte Carlo Tree Search for Code Generation
by: Lin, Yuanyuan, et al.
Published: (2025)
by: Lin, Yuanyuan, et al.
Published: (2025)
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
by: Zheng, Xiang, et al.
Published: (2026)
by: Zheng, Xiang, et al.
Published: (2026)
Graph Ranking Contrastive Learning: A Extremely Simple yet Efficient Method
by: Hu, Yulan, et al.
Published: (2023)
by: Hu, Yulan, et al.
Published: (2023)
Rewarding Structural Conformance of Reasoning using Process Mining
by: Lee, Yongjae, et al.
Published: (2025)
by: Lee, Yongjae, et al.
Published: (2025)
Process Reward Agents for Steering Knowledge-Intensive Reasoning
by: Sohn, Jiwoong, et al.
Published: (2026)
by: Sohn, Jiwoong, et al.
Published: (2026)
CFMS: A Coarse-to-Fine Multimodal Synthesis Framework for Enhanced Tabular Reasoning
by: Huang, Qixian, et al.
Published: (2026)
by: Huang, Qixian, et al.
Published: (2026)
Optimizing Agentic Reasoning with Retrieval via Synthetic Semantic Information Gain Reward
by: Hu, Senkang, et al.
Published: (2026)
by: Hu, Senkang, et al.
Published: (2026)
Not All Preferences Are Created Equal: Stability-Aware and Gradient-Efficient Alignment for Reasoning Models
by: Wu, Hui, et al.
Published: (2026)
by: Wu, Hui, et al.
Published: (2026)
Similar Items
-
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
by: Ouyang, Sheng, et al.
Published: (2025) -
Towards Comprehensive Preference Data Collection for Reward Modeling
by: Hu, Yulan, et al.
Published: (2024) -
GUNDAM: Aligning Large Language Models with Graph Understanding
by: Ouyang, Sheng, et al.
Published: (2024) -
Refining Latent Representations: A Generative SSL Approach for Heterogeneous Graph Learning
by: Hu, Yulan, et al.
Published: (2023) -
VIGraph: Generative Self-supervised Learning for Class-Imbalanced Node Classification
by: Hu, Yulan, et al.
Published: (2023)