DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yuanhao, Song, Juntong, Zhang, Hanning, Zhang, Tong, Niu, Cheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenGenAlign: A Preference Dataset and Benchmark for Trustworthy Reward Modeling in Open-Ended, Long-Context Generation
by: Zhang, Hanning, et al.
Published: (2025)
by: Zhang, Hanning, et al.
Published: (2025)
Enhancing Mathematical Reasoning in LLMs by Stepwise Correction
by: Wu, Zhenyu, et al.
Published: (2024)
by: Wu, Zhenyu, et al.
Published: (2024)
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
by: Niu, Cheng, et al.
Published: (2023)
by: Niu, Cheng, et al.
Published: (2023)
VeraCT Scan: Retrieval-Augmented Fake News Detection with Justifiable Reasoning
by: Niu, Cheng, et al.
Published: (2024)
by: Niu, Cheng, et al.
Published: (2024)
Enhancing Dialogue State Tracking Models through LLM-backed User-Agents Simulation
by: Niu, Cheng, et al.
Published: (2024)
by: Niu, Cheng, et al.
Published: (2024)
Step Potential Advantage Estimation: Harnessing Intermediate Confidence and Correctness for Efficient Mathematical Reasoning
by: Wu, Fei, et al.
Published: (2026)
by: Wu, Fei, et al.
Published: (2026)
Reward-SQL: Boosting Text-to-SQL via Stepwise Reasoning and Process-Supervised Rewards
by: Zhang, Yuxin, et al.
Published: (2025)
by: Zhang, Yuxin, et al.
Published: (2025)
SGR: A Stepwise Reasoning Framework for LLMs with External Subgraph Generation
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards
by: Yun, Jaehoon, et al.
Published: (2025)
by: Yun, Jaehoon, et al.
Published: (2025)
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
by: Zhang, Jianghangfan, et al.
Published: (2025)
by: Zhang, Jianghangfan, et al.
Published: (2025)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
by: Zhang, Zhenru, et al.
Published: (2025)
by: Zhang, Zhenru, et al.
Published: (2025)
StepWiser: Stepwise Generative Judges for Wiser Reasoning
by: Xiong, Wei, et al.
Published: (2025)
by: Xiong, Wei, et al.
Published: (2025)
Entropy-Regularized Process Reward Model
by: Zhang, Hanning, et al.
Published: (2024)
by: Zhang, Hanning, et al.
Published: (2024)
A Stepwise-Enhanced Reasoning Framework for Large Language Models Based on External Subgraph Generation
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Stepwise Self-Consistent Mathematical Reasoning with Large Language Models
by: Zhao, Zilong, et al.
Published: (2024)
by: Zhao, Zilong, et al.
Published: (2024)
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
by: Lyu, Chengqi, et al.
Published: (2025)
by: Lyu, Chengqi, et al.
Published: (2025)
Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning
by: Xu, Huimin, et al.
Published: (2025)
by: Xu, Huimin, et al.
Published: (2025)
Towards Stepwise Domain Knowledge-Driven Reasoning Optimization and Reflection Improvement
by: Liu, Chengyuan, et al.
Published: (2025)
by: Liu, Chengyuan, et al.
Published: (2025)
Prioritizing the Best: Incentivizing Reliable Multimodal Reasoning by Rewarding Beyond Answer Correctness
by: Jia, Mengzhao, et al.
Published: (2026)
by: Jia, Mengzhao, et al.
Published: (2026)
Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
by: Yuan, Youliang, et al.
Published: (2025)
by: Yuan, Youliang, et al.
Published: (2025)
RM-R1: Reward Modeling as Reasoning
by: Chen, Xiusi, et al.
Published: (2025)
by: Chen, Xiusi, et al.
Published: (2025)
Rubric-Guided Process Reward for Stepwise Model Routing
by: Ye, Shenghao, et al.
Published: (2026)
by: Ye, Shenghao, et al.
Published: (2026)
Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
by: Xu, Zhichao, et al.
Published: (2025)
by: Xu, Zhichao, et al.
Published: (2025)
SMRC: Aligning Large Language Models with Student Reasoning for Mathematical Error Correction
by: Zeng, Biaojie, et al.
Published: (2025)
by: Zeng, Biaojie, et al.
Published: (2025)
Evaluating Robustness of Reward Models for Mathematical Reasoning
by: Kim, Sunghwan, et al.
Published: (2024)
by: Kim, Sunghwan, et al.
Published: (2024)
ETR: Entropy Trend Reward for Efficient Chain-of-Thought Reasoning
by: Xiong, Xuan, et al.
Published: (2026)
by: Xiong, Xuan, et al.
Published: (2026)
An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning
by: Sun, Wei, et al.
Published: (2025)
by: Sun, Wei, et al.
Published: (2025)
Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning
by: Rajaee, Sara, et al.
Published: (2025)
by: Rajaee, Sara, et al.
Published: (2025)
Stepwise Penalization for Length-Efficient Chain-of-Thought Reasoning
by: Li, Xintong, et al.
Published: (2026)
by: Li, Xintong, et al.
Published: (2026)
Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
by: Hwang, Hyeonbin, et al.
Published: (2024)
by: Hwang, Hyeonbin, et al.
Published: (2024)
Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive Thinking
by: Huang, Weiyang, et al.
Published: (2026)
by: Huang, Weiyang, et al.
Published: (2026)
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning
by: Zhao, Jun, et al.
Published: (2024)
by: Zhao, Jun, et al.
Published: (2024)
Stepwise Informativeness Search for Efficient and Effective LLM Reasoning
by: Wang, Siyuan, et al.
Published: (2025)
by: Wang, Siyuan, et al.
Published: (2025)
Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
by: Yao, Jiarui, et al.
Published: (2025)
by: Yao, Jiarui, et al.
Published: (2025)
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Dancing with Critiques: Enhancing LLM Reasoning with Stepwise Natural Language Self-Critique
by: Li, Yansi, et al.
Published: (2025)
by: Li, Yansi, et al.
Published: (2025)
Triviality Corrected Endogenous Reward
by: Wang, Xinda, et al.
Published: (2026)
by: Wang, Xinda, et al.
Published: (2026)
THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
by: Chang, Qikai, et al.
Published: (2025)
by: Chang, Qikai, et al.
Published: (2025)
Similar Items
-
OpenGenAlign: A Preference Dataset and Benchmark for Trustworthy Reward Modeling in Open-Ended, Long-Context Generation
by: Zhang, Hanning, et al.
Published: (2025) -
Enhancing Mathematical Reasoning in LLMs by Stepwise Correction
by: Wu, Zhenyu, et al.
Published: (2024) -
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
by: Niu, Cheng, et al.
Published: (2023) -
VeraCT Scan: Retrieval-Augmented Fake News Detection with Justifiable Reasoning
by: Niu, Cheng, et al.
Published: (2024) -
Enhancing Dialogue State Tracking Models through LLM-backed User-Agents Simulation
by: Niu, Cheng, et al.
Published: (2024)