Rewarding Graph Reasoning Process makes LLMs more Generalized Reasoners
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Peng, Miao, Chen, Nuo, Suo, Zongrui, Li, Jia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problems
von: Wang, Yuyao, et al.
Veröffentlicht: (2025)
von: Wang, Yuyao, et al.
Veröffentlicht: (2025)
ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
von: Fan, Lishui, et al.
Veröffentlicht: (2025)
von: Fan, Lishui, et al.
Veröffentlicht: (2025)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
von: Zhang, Zhenru, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenru, et al.
Veröffentlicht: (2025)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
von: Luo, Ruilin, et al.
Veröffentlicht: (2025)
von: Luo, Ruilin, et al.
Veröffentlicht: (2025)
What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
von: Lai, Xin, et al.
Veröffentlicht: (2024)
von: Lai, Xin, et al.
Veröffentlicht: (2024)
Deja vu: Contrastive Historical Modeling with Prefix-tuning for Temporal Knowledge Graph Reasoning
von: Peng, Miao, et al.
Veröffentlicht: (2024)
von: Peng, Miao, et al.
Veröffentlicht: (2024)
Incentivizing In-depth Reasoning over Long Contexts with Process Advantage Shaping
von: Peng, Miao, et al.
Veröffentlicht: (2026)
von: Peng, Miao, et al.
Veröffentlicht: (2026)
RM-R1: Reward Modeling as Reasoning
von: Chen, Xiusi, et al.
Veröffentlicht: (2025)
von: Chen, Xiusi, et al.
Veröffentlicht: (2025)
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
ControlMath: Controllable Data Generation Promotes Math Generalist Models
von: Chen, Nuo, et al.
Veröffentlicht: (2024)
von: Chen, Nuo, et al.
Veröffentlicht: (2024)
FlowRL: Matching Reward Distributions for LLM Reasoning
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
Evaluating Robustness of Reward Models for Mathematical Reasoning
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
von: Zhao, Guojiang, et al.
Veröffentlicht: (2025)
von: Zhao, Guojiang, et al.
Veröffentlicht: (2025)
DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological Reasoning
von: Chan, Chi-Min, et al.
Veröffentlicht: (2026)
von: Chan, Chi-Min, et al.
Veröffentlicht: (2026)
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training
von: Bai, Yuyang, et al.
Veröffentlicht: (2026)
von: Bai, Yuyang, et al.
Veröffentlicht: (2026)
Can Post-Training Transform LLMs into Causal Reasoners?
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
Verifier-Backed Hard Problem Generation for Mathematical Reasoning
von: Lai, Yuhang, et al.
Veröffentlicht: (2026)
von: Lai, Yuhang, et al.
Veröffentlicht: (2026)
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
von: Yu, Erxin, et al.
Veröffentlicht: (2025)
von: Yu, Erxin, et al.
Veröffentlicht: (2025)
Chain of Execution Supervision Promotes General Reasoning in Large Language Models
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
Reverse Thinking Makes LLMs Stronger Reasoners
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2024)
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2024)
ReCellTy: Domain-Specific Knowledge Graph Retrieval-Augmented LLMs Reasoning Workflow for Single-Cell Annotation
von: Han, Dezheng, et al.
Veröffentlicht: (2025)
von: Han, Dezheng, et al.
Veröffentlicht: (2025)
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization
von: Kawakami, Wataru, et al.
Veröffentlicht: (2025)
von: Kawakami, Wataru, et al.
Veröffentlicht: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
von: Chegini, Atoosa, et al.
Veröffentlicht: (2026)
von: Chegini, Atoosa, et al.
Veröffentlicht: (2026)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
von: Damani, Mehul, et al.
Veröffentlicht: (2025)
von: Damani, Mehul, et al.
Veröffentlicht: (2025)
Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation
von: Padarha, Shreyansh
Veröffentlicht: (2025)
von: Padarha, Shreyansh
Veröffentlicht: (2025)
REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
von: Stojanovski, Zafir, et al.
Veröffentlicht: (2025)
von: Stojanovski, Zafir, et al.
Veröffentlicht: (2025)
On Designing Effective RL Reward at Training Time for LLM Reasoning
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2024)
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2024)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025)
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025)
More Bang for the Buck: Process Reward Modeling with Entropy-Driven Uncertainty
von: Cao, Lang, et al.
Veröffentlicht: (2025)
von: Cao, Lang, et al.
Veröffentlicht: (2025)
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
von: Suvarna, Ashima, et al.
Veröffentlicht: (2026)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2026)
Graph Neural Network-Based Entity Extraction and Relationship Reasoning in Complex Knowledge Graphs
von: Du, Junliang, et al.
Veröffentlicht: (2024)
von: Du, Junliang, et al.
Veröffentlicht: (2024)
Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs
von: Kim, Jaemin, et al.
Veröffentlicht: (2025)
von: Kim, Jaemin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problems
von: Wang, Yuyao, et al.
Veröffentlicht: (2025) -
ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
von: Fan, Lishui, et al.
Veröffentlicht: (2025) -
The Lessons of Developing Process Reward Models in Mathematical Reasoning
von: Zhang, Zhenru, et al.
Veröffentlicht: (2025) -
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
von: Luo, Ruilin, et al.
Veröffentlicht: (2025) -
What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding
von: Li, Ming, et al.
Veröffentlicht: (2025)