GPO: Learning from Critical Steps to Improve LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Jiahao, Cheng, Zelei, Wu, Xian, Xing, Xinyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization
by: Yu, Jiahao, et al.
Published: (2025)
by: Yu, Jiahao, et al.
Published: (2025)
A Survey on Explainable Deep Reinforcement Learning
by: Cheng, Zelei, et al.
Published: (2025)
by: Cheng, Zelei, et al.
Published: (2025)
RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with Explanation
by: Cheng, Zelei, et al.
Published: (2024)
by: Cheng, Zelei, et al.
Published: (2024)
Soft-Label Integration for Robust Toxicity Classification
by: Cheng, Zelei, et al.
Published: (2024)
by: Cheng, Zelei, et al.
Published: (2024)
BlockScan: Detecting Anomalies in Blockchain Transactions
by: Yu, Jiahao, et al.
Published: (2024)
by: Yu, Jiahao, et al.
Published: (2024)
CPL: Critical Plan Step Learning Boosts LLM Generalization in Reasoning Tasks
by: Wang, Tianlong, et al.
Published: (2024)
by: Wang, Tianlong, et al.
Published: (2024)
Graph-Augmented Reasoning: Evolving Step-by-Step Knowledge Graph Retrieval for LLM Reasoning
by: Wu, Wenjie, et al.
Published: (2025)
by: Wu, Wenjie, et al.
Published: (2025)
UC-MOA: Utility-Conditioned Multi-Objective Alignment for Distributional Pareto-Optimality
by: Cheng, Zelei, et al.
Published: (2025)
by: Cheng, Zelei, et al.
Published: (2025)
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
by: Yu, Jiahao, et al.
Published: (2023)
by: Yu, Jiahao, et al.
Published: (2023)
Pre-Act: Multi-Step Planning and Reasoning Improves Acting in LLM Agents
by: Rawat, Mrinal, et al.
Published: (2025)
by: Rawat, Mrinal, et al.
Published: (2025)
Offline Reinforcement Learning for LLM Multi-Step Reasoning
by: Wang, Huaijie, et al.
Published: (2024)
by: Wang, Huaijie, et al.
Published: (2024)
PhGPO: Pheromone-Guided Policy Optimization for Long-Horizon Tool Planning
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step Reasoning
by: Cao, Lang, et al.
Published: (2024)
by: Cao, Lang, et al.
Published: (2024)
A Neuro-Symbolic Framework for Reasoning under Perceptual Uncertainty: Bridging Continuous Perception and Discrete Symbolic Planning
by: Wu, Jiahao, et al.
Published: (2025)
by: Wu, Jiahao, et al.
Published: (2025)
rePIRL: Learn PRM with Inverse RL for LLM Reasoning
by: Wu, Xian, et al.
Published: (2026)
by: Wu, Xian, et al.
Published: (2026)
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
Adaptive Selection of Symbolic Languages for Improving LLM Logical Reasoning
by: Wang, Xiangyu, et al.
Published: (2025)
by: Wang, Xiangyu, et al.
Published: (2025)
Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM Agents
by: Zhen, Shuai, et al.
Published: (2026)
by: Zhen, Shuai, et al.
Published: (2026)
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
by: Hao, Shibo, et al.
Published: (2024)
by: Hao, Shibo, et al.
Published: (2024)
Asking the Right Questions: Improving Reasoning with Generated Stepping Stones
by: Hu, Hengyuan, et al.
Published: (2026)
by: Hu, Hengyuan, et al.
Published: (2026)
Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
by: Wang, Yiding, et al.
Published: (2025)
by: Wang, Yiding, et al.
Published: (2025)
Scientific Logicality Enriched Methodology for LLM Reasoning: A Practice in Physics
by: Yu, Zhaoxin, et al.
Published: (2026)
by: Yu, Zhaoxin, et al.
Published: (2026)
R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning
by: Li, Yuan, et al.
Published: (2025)
by: Li, Yuan, et al.
Published: (2025)
SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
by: Xu, Yuyang, et al.
Published: (2025)
by: Xu, Yuyang, et al.
Published: (2025)
Assessing Prompt Injection Risks in 200+ Custom GPTs
by: Yu, Jiahao, et al.
Published: (2023)
by: Yu, Jiahao, et al.
Published: (2023)
Interactive Learning for LLM Reasoning
by: Lin, Hehai, et al.
Published: (2025)
by: Lin, Hehai, et al.
Published: (2025)
LLM Reasoning with Process Rewards for Outcome-Guided Steps
by: Rezaei, Mohammad, et al.
Published: (2026)
by: Rezaei, Mohammad, et al.
Published: (2026)
On the Step Length Confounding in LLM Reasoning Data Selection
by: Wang, Bing, et al.
Published: (2026)
by: Wang, Bing, et al.
Published: (2026)
Reinforcement Learning Improves LLM Accuracy and Reasoning in Disease Classification from Radiology Reports
by: Wei, Yishu, et al.
Published: (2026)
by: Wei, Yishu, et al.
Published: (2026)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
by: Saha, Swarnadeep, et al.
Published: (2025)
by: Saha, Swarnadeep, et al.
Published: (2025)
Masked Thought: Simply Masking Partial Reasoning Steps Can Improve Mathematical Reasoning Learning of Language Models
by: Chen, Changyu, et al.
Published: (2024)
by: Chen, Changyu, et al.
Published: (2024)
On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents
by: Zou, Deyu, et al.
Published: (2026)
by: Zou, Deyu, et al.
Published: (2026)
Two-Stage Reasoning-Infused Learning: Improving Classification with LLM-Generated Reasoning
by: Henrichsen, Mads, et al.
Published: (2025)
by: Henrichsen, Mads, et al.
Published: (2025)
ProCeedRL: Process Critic with Exploratory Demonstration Reinforcement Learning for LLM Agentic Reasoning
by: Gao, Jingyue, et al.
Published: (2026)
by: Gao, Jingyue, et al.
Published: (2026)
Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
by: Xiong, Weimin, et al.
Published: (2024)
by: Xiong, Weimin, et al.
Published: (2024)
Benchmarking Multi-Step Legal Reasoning and Analyzing Chain-of-Thought Effects in Large Language Models
by: Yu, Wenhan, et al.
Published: (2025)
by: Yu, Wenhan, et al.
Published: (2025)
Discovering Process-Outcome Credit in Multi-Step LLM Reasoning
by: Wang, Xiangwei, et al.
Published: (2026)
by: Wang, Xiangwei, et al.
Published: (2026)
Beyond Meta-Reasoning: Metacognitive Consolidation for Self-Improving LLM Reasoning
by: Zhuang, Ziqing, et al.
Published: (2026)
by: Zhuang, Ziqing, et al.
Published: (2026)
Experiential Reflective Learning for Self-Improving LLM Agents
by: Allard, Marc-Antoine, et al.
Published: (2026)
by: Allard, Marc-Antoine, et al.
Published: (2026)
ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
Similar Items
-
Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization
by: Yu, Jiahao, et al.
Published: (2025) -
A Survey on Explainable Deep Reinforcement Learning
by: Cheng, Zelei, et al.
Published: (2025) -
RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with Explanation
by: Cheng, Zelei, et al.
Published: (2024) -
Soft-Label Integration for Robust Toxicity Classification
by: Cheng, Zelei, et al.
Published: (2024) -
BlockScan: Detecting Anomalies in Blockchain Transactions
by: Yu, Jiahao, et al.
Published: (2024)