GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models
Fuente:
arXiv
Guardado en:
| Autor principal: | Wang, Zhijie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning
por: Yang, Zhaohui, et al.
Publicado: (2025)
por: Yang, Zhaohui, et al.
Publicado: (2025)
GRPO is Secretly a Process Reward Model
por: Sullivan, Michael, et al.
Publicado: (2025)
por: Sullivan, Michael, et al.
Publicado: (2025)
Large Language Models and Mathematical Reasoning Failures
por: Boye, Johan, et al.
Publicado: (2025)
por: Boye, Johan, et al.
Publicado: (2025)
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
por: Hu, Yulan, et al.
Publicado: (2025)
por: Hu, Yulan, et al.
Publicado: (2025)
Mathematical Computation and Reasoning Errors by Large Language Models
por: Zhang, Liang, et al.
Publicado: (2025)
por: Zhang, Liang, et al.
Publicado: (2025)
A Survey on Large Language Models for Mathematical Reasoning
por: Wang, Peng-Yuan, et al.
Publicado: (2025)
por: Wang, Peng-Yuan, et al.
Publicado: (2025)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
por: Xiao, Wenyi, et al.
Publicado: (2025)
por: Xiao, Wenyi, et al.
Publicado: (2025)
Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning
por: Zhu, Jiachen, et al.
Publicado: (2025)
por: Zhu, Jiachen, et al.
Publicado: (2025)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
por: Xu, Yuanda, et al.
Publicado: (2026)
por: Xu, Yuanda, et al.
Publicado: (2026)
Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
por: Lu, Leo, et al.
Publicado: (2025)
por: Lu, Leo, et al.
Publicado: (2025)
A Survey on Mathematical Reasoning and Optimization with Large Language Models
por: Forootani, Ali
Publicado: (2025)
por: Forootani, Ali
Publicado: (2025)
Evaluating Robustness of Reward Models for Mathematical Reasoning
por: Kim, Sunghwan, et al.
Publicado: (2024)
por: Kim, Sunghwan, et al.
Publicado: (2024)
M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models
por: Wang, Junjian, et al.
Publicado: (2026)
por: Wang, Junjian, et al.
Publicado: (2026)
Teaching Large Reasoning Models Effective Reflection
por: Wang, Hanbin, et al.
Publicado: (2026)
por: Wang, Hanbin, et al.
Publicado: (2026)
Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation
por: Dai, Yanqi, et al.
Publicado: (2026)
por: Dai, Yanqi, et al.
Publicado: (2026)
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
por: Dipta, Shubhashis Roy, et al.
Publicado: (2026)
por: Dipta, Shubhashis Roy, et al.
Publicado: (2026)
Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning
por: Rajaee, Sara, et al.
Publicado: (2025)
por: Rajaee, Sara, et al.
Publicado: (2025)
Shaping Explanations: Semantic Reward Modeling with Encoder-Only Transformers for GRPO
por: Pappone, Francesco, et al.
Publicado: (2025)
por: Pappone, Francesco, et al.
Publicado: (2025)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
por: Luo, Ruilin, et al.
Publicado: (2025)
por: Luo, Ruilin, et al.
Publicado: (2025)
Key-Point-Driven Mathematical Reasoning Distillation of Large Language Model
por: Zhu, Xunyu, et al.
Publicado: (2024)
por: Zhu, Xunyu, et al.
Publicado: (2024)
Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning
por: He, Qianxi, et al.
Publicado: (2025)
por: He, Qianxi, et al.
Publicado: (2025)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
por: Zhang, Zhenru, et al.
Publicado: (2025)
por: Zhang, Zhenru, et al.
Publicado: (2025)
Efficient Paths and Dense Rewards: Probabilistic Flow Reasoning for Large Language Models
por: Liu, Yan, et al.
Publicado: (2026)
por: Liu, Yan, et al.
Publicado: (2026)
ShapE-GRPO: Shapley-Enhanced Reward Allocation for Multi-Candidate LLM Training
por: Ai, Rui, et al.
Publicado: (2026)
por: Ai, Rui, et al.
Publicado: (2026)
AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
por: Huang, Runhui, et al.
Publicado: (2026)
por: Huang, Runhui, et al.
Publicado: (2026)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
por: Tan, Hongze, et al.
Publicado: (2025)
por: Tan, Hongze, et al.
Publicado: (2025)
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning
por: Zhao, Jun, et al.
Publicado: (2024)
por: Zhao, Jun, et al.
Publicado: (2024)
Numerical Sensitivity and Robustness: Exploring the Flaws of Mathematical Reasoning in Large Language Models
por: Sun, Zhishen, et al.
Publicado: (2025)
por: Sun, Zhishen, et al.
Publicado: (2025)
Revisiting Anthropomorphic Reflection Markers in Large Language Model Reasoning
por: Yu, Yahan, et al.
Publicado: (2026)
por: Yu, Yahan, et al.
Publicado: (2026)
Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning
por: Chen, Benteng, et al.
Publicado: (2026)
por: Chen, Benteng, et al.
Publicado: (2026)
From Reasoning to Code: GRPO Optimization for Underrepresented Languages
por: Pennino, Federico, et al.
Publicado: (2025)
por: Pennino, Federico, et al.
Publicado: (2025)
MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
por: Chen, Jinhao, et al.
Publicado: (2025)
por: Chen, Jinhao, et al.
Publicado: (2025)
What Are Step-Level Reward Models Rewarding? Counterintuitive Findings from MCTS-Boosted Mathematical Reasoning
por: Ma, Yiran, et al.
Publicado: (2024)
por: Ma, Yiran, et al.
Publicado: (2024)
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
por: Mansouri, Omar El, et al.
Publicado: (2025)
por: Mansouri, Omar El, et al.
Publicado: (2025)
Towards Hierarchical Multi-Step Reward Models for Enhanced Reasoning in Large Language Models
por: Wang, Teng, et al.
Publicado: (2025)
por: Wang, Teng, et al.
Publicado: (2025)
Enhancing Mathematical Reasoning in Large Language Models with Self-Consistency-Based Hallucination Detection
por: Liu, MingShan, et al.
Publicado: (2025)
por: Liu, MingShan, et al.
Publicado: (2025)
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
por: Zhang, Xiaoying, et al.
Publicado: (2025)
por: Zhang, Xiaoying, et al.
Publicado: (2025)
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
por: Mirzadeh, Iman, et al.
Publicado: (2024)
por: Mirzadeh, Iman, et al.
Publicado: (2024)
CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal Knowledge
por: Zan, Lei, et al.
Publicado: (2025)
por: Zan, Lei, et al.
Publicado: (2025)
Systematic Optimization of Open Source Large Language Models for Mathematical Reasoning
por: Pawar, Pranav, et al.
Publicado: (2025)
por: Pawar, Pranav, et al.
Publicado: (2025)
Ejemplares similares
-
Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning
por: Yang, Zhaohui, et al.
Publicado: (2025) -
GRPO is Secretly a Process Reward Model
por: Sullivan, Michael, et al.
Publicado: (2025) -
Large Language Models and Mathematical Reasoning Failures
por: Boye, Johan, et al.
Publicado: (2025) -
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
por: Hu, Yulan, et al.
Publicado: (2025) -
Mathematical Computation and Reasoning Errors by Large Language Models
por: Zhang, Liang, et al.
Publicado: (2025)