ProCeedRL: Process Critic with Exploratory Demonstration Reinforcement Learning for LLM Agentic Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Jingyue, Guo, Yanjiang, Chen, Xiaoshuai, Chen, Jianyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MARGE: Improving Math Reasoning for LLMs with Guided Exploration
by: Gao, Jingyue, et al.
Published: (2025)
by: Gao, Jingyue, et al.
Published: (2025)
Prediction with Action: Visual Policy Learning via Joint Denoising Process
by: Guo, Yanjiang, et al.
Published: (2024)
by: Guo, Yanjiang, et al.
Published: (2024)
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
by: Guo, Yanjiang, et al.
Published: (2025)
by: Guo, Yanjiang, et al.
Published: (2025)
DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
by: Guo, Yanjiang, et al.
Published: (2023)
by: Guo, Yanjiang, et al.
Published: (2023)
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
by: Liu, Mingjie, et al.
Published: (2025)
by: Liu, Mingjie, et al.
Published: (2025)
Efficient and Generalized end-to-end Autonomous Driving System with Latent Deep Reinforcement Learning and Demonstrations
by: Tang, Zuojin, et al.
Published: (2024)
by: Tang, Zuojin, et al.
Published: (2024)
Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
by: Xie, Tian, et al.
Published: (2025)
by: Xie, Tian, et al.
Published: (2025)
RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning
by: Hao, Qianyue, et al.
Published: (2025)
by: Hao, Qianyue, et al.
Published: (2025)
MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism
by: Liu, Shulin, et al.
Published: (2025)
by: Liu, Shulin, et al.
Published: (2025)
UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
by: Zhang, Jianke, et al.
Published: (2025)
by: Zhang, Jianke, et al.
Published: (2025)
Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering
by: Li, Jingyue, et al.
Published: (2026)
by: Li, Jingyue, et al.
Published: (2026)
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
by: Li, Wanli, et al.
Published: (2026)
by: Li, Wanli, et al.
Published: (2026)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
by: Chen, Wanyi, et al.
Published: (2026)
by: Chen, Wanyi, et al.
Published: (2026)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
by: Xu, Ran, et al.
Published: (2025)
by: Xu, Ran, et al.
Published: (2025)
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation
by: Hou, Hongru, et al.
Published: (2026)
by: Hou, Hongru, et al.
Published: (2026)
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
by: Brantley, Kianté, et al.
Published: (2025)
by: Brantley, Kianté, et al.
Published: (2025)
KnowRL: Boosting LLM Reasoning via Reinforcement Learning with Minimal-Sufficient Knowledge Guidance
by: Yu, Linhao, et al.
Published: (2026)
by: Yu, Linhao, et al.
Published: (2026)
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
by: Zhang, Jianke, et al.
Published: (2025)
by: Zhang, Jianke, et al.
Published: (2025)
rePIRL: Learn PRM with Inverse RL for LLM Reasoning
by: Wu, Xian, et al.
Published: (2026)
by: Wu, Xian, et al.
Published: (2026)
Verifiable Process Rewards for Agentic Reasoning
by: Yuan, Huining, et al.
Published: (2026)
by: Yuan, Huining, et al.
Published: (2026)
ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
by: Zhang, Hao, et al.
Published: (2026)
by: Zhang, Hao, et al.
Published: (2026)
Role-RL: Online Long-Context Processing with Role Reinforcement Learning for Distinct LLMs in Their Optimal Roles
by: He, Lewei, et al.
Published: (2024)
by: He, Lewei, et al.
Published: (2024)
Learning Reasoning Rewards from Expert Demonstrations with Inverse Reinforcement Learning
by: Fanconi, Claudio, et al.
Published: (2025)
by: Fanconi, Claudio, et al.
Published: (2025)
MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
by: Xu, Yifan, et al.
Published: (2025)
by: Xu, Yifan, et al.
Published: (2025)
Cog-Rethinker: Hierarchical Metacognitive Reinforcement Learning for LLM Reasoning
by: Sun, Zexu, et al.
Published: (2025)
by: Sun, Zexu, et al.
Published: (2025)
Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning
by: Gu, Xinyang, et al.
Published: (2024)
by: Gu, Xinyang, et al.
Published: (2024)
AgentReputation: A Decentralized Agentic AI Reputation Framework
by: Chishti, Mohd Sameen, et al.
Published: (2026)
by: Chishti, Mohd Sameen, et al.
Published: (2026)
TKG-Thinker: Towards Dynamic Reasoning over Temporal Knowledge Graphs via Agentic Reinforcement Learning
by: Jiang, Zihao, et al.
Published: (2026)
by: Jiang, Zihao, et al.
Published: (2026)
SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning
by: He, Jiashu, et al.
Published: (2025)
by: He, Jiashu, et al.
Published: (2025)
Process In-Context Learning: Enhancing Mathematical Reasoning via Dynamic Demonstration Insertion
by: Gao, Ang, et al.
Published: (2026)
by: Gao, Ang, et al.
Published: (2026)
AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
by: Zhang, Hanchen, et al.
Published: (2025)
by: Zhang, Hanchen, et al.
Published: (2025)
CARL: Criticality-Aware Agentic Reinforcement Learning
by: Shen, Leyang, et al.
Published: (2025)
by: Shen, Leyang, et al.
Published: (2025)
SuperRL: Reinforcement Learning with Supervision to Boost Language Model Reasoning
by: Liu, Yihao, et al.
Published: (2025)
by: Liu, Yihao, et al.
Published: (2025)
Scheduling Your LLM Reinforcement Learning with Reasoning Trees
by: Wang, Hong, et al.
Published: (2025)
by: Wang, Hong, et al.
Published: (2025)
Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
by: Singh, Joykirat, et al.
Published: (2025)
by: Singh, Joykirat, et al.
Published: (2025)
Exploratory Diffusion Model for Unsupervised Reinforcement Learning
by: Ying, Chengyang, et al.
Published: (2025)
by: Ying, Chengyang, et al.
Published: (2025)
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
Coupled Variational Reinforcement Learning for Language Model General Reasoning
by: Wen, Xueru, et al.
Published: (2025)
by: Wen, Xueru, et al.
Published: (2025)
ProMind-LLM: Proactive Mental Health Care via Causal Reasoning with Sensor Data
by: Zheng, Xinzhe, et al.
Published: (2025)
by: Zheng, Xinzhe, et al.
Published: (2025)
Similar Items
-
MARGE: Improving Math Reasoning for LLMs with Guided Exploration
by: Gao, Jingyue, et al.
Published: (2025) -
Prediction with Action: Visual Policy Learning via Joint Denoising Process
by: Guo, Yanjiang, et al.
Published: (2024) -
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
by: Guo, Yanjiang, et al.
Published: (2025) -
DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
by: Guo, Yanjiang, et al.
Published: (2023) -
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
by: Liu, Mingjie, et al.
Published: (2025)