Learning from the Irrecoverable: Error-Localized Policy Optimization for Tool-Integrated LLM Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Qiao, Zhu, Yuke, Ge, Chao, Yang, Lei, Shen, Ying, Zheng, Bo, Guo, Sheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated Penalty
von: Yu, Zewei, et al.
Veröffentlicht: (2026)
von: Yu, Zewei, et al.
Veröffentlicht: (2026)
PORTool: Importance-Aware Policy Optimization with Rewarded Tree for Multi-Tool-Integrated Reasoning
von: Wu, Feijie, et al.
Veröffentlicht: (2025)
von: Wu, Feijie, et al.
Veröffentlicht: (2025)
AdaTIR: Adaptive Tool-Integrated Reasoning via Difficulty-Aware Policy Optimization
von: Fang, Zhaiyu, et al.
Veröffentlicht: (2026)
von: Fang, Zhaiyu, et al.
Veröffentlicht: (2026)
Multi-Agent Tool-Integrated Policy Optimization
von: Mo, Zhanfeng, et al.
Veröffentlicht: (2025)
von: Mo, Zhanfeng, et al.
Veröffentlicht: (2025)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
von: Xu, Ran, et al.
Veröffentlicht: (2025)
von: Xu, Ran, et al.
Veröffentlicht: (2025)
AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning
von: Ma, Qinghe, et al.
Veröffentlicht: (2026)
von: Ma, Qinghe, et al.
Veröffentlicht: (2026)
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
von: Ding, Yifeng, et al.
Veröffentlicht: (2025)
von: Ding, Yifeng, et al.
Veröffentlicht: (2025)
Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
von: Zhao, Yufeng, et al.
Veröffentlicht: (2025)
von: Zhao, Yufeng, et al.
Veröffentlicht: (2025)
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning
von: Cheng, Qianjia, et al.
Veröffentlicht: (2026)
von: Cheng, Qianjia, et al.
Veröffentlicht: (2026)
Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient
von: Li, Changming, et al.
Veröffentlicht: (2026)
von: Li, Changming, et al.
Veröffentlicht: (2026)
Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning
von: He, Yinhan, et al.
Veröffentlicht: (2026)
von: He, Yinhan, et al.
Veröffentlicht: (2026)
Temporal Consistency for LLM Reasoning Process Error Identification
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
von: Wei, Jiaqi, et al.
Veröffentlicht: (2026)
von: Wei, Jiaqi, et al.
Veröffentlicht: (2026)
Guiding LLM-based Loop Invariant Synthesis via Feedback on Local Reasoning Errors
von: Li, Tianchi, et al.
Veröffentlicht: (2026)
von: Li, Tianchi, et al.
Veröffentlicht: (2026)
LLM Agents Already Know When to Call Tools -- Even Without Reasoning
von: Sun, Chung-En, et al.
Veröffentlicht: (2026)
von: Sun, Chung-En, et al.
Veröffentlicht: (2026)
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
von: He, Yancheng, et al.
Veröffentlicht: (2025)
von: He, Yancheng, et al.
Veröffentlicht: (2025)
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
von: Deng, Jingcheng, et al.
Veröffentlicht: (2026)
von: Deng, Jingcheng, et al.
Veröffentlicht: (2026)
Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoning
von: Wang, Ziyan, et al.
Veröffentlicht: (2025)
von: Wang, Ziyan, et al.
Veröffentlicht: (2025)
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
Discovery and Reinforcement of Tool-Integrated Reasoning Chains via Rollout Trees
von: Li, Kun, et al.
Veröffentlicht: (2026)
von: Li, Kun, et al.
Veröffentlicht: (2026)
Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning
von: Zhang, Dongxu, et al.
Veröffentlicht: (2025)
von: Zhang, Dongxu, et al.
Veröffentlicht: (2025)
THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
von: Chang, Qikai, et al.
Veröffentlicht: (2025)
von: Chang, Qikai, et al.
Veröffentlicht: (2025)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use
von: Pang, Renning, et al.
Veröffentlicht: (2026)
von: Pang, Renning, et al.
Veröffentlicht: (2026)
Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
von: Wu, Junde, et al.
Veröffentlicht: (2025)
von: Wu, Junde, et al.
Veröffentlicht: (2025)
AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning
von: Wei, Yifan, et al.
Veröffentlicht: (2025)
von: Wei, Yifan, et al.
Veröffentlicht: (2025)
LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning
von: Wang, Jianing, et al.
Veröffentlicht: (2026)
von: Wang, Jianing, et al.
Veröffentlicht: (2026)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
von: Gou, Zhibin, et al.
Veröffentlicht: (2023)
von: Gou, Zhibin, et al.
Veröffentlicht: (2023)
Error as a Lens: Probing LLM Reasoning through Synthetic Misconception Generation
von: Yang, Xinming, et al.
Veröffentlicht: (2026)
von: Yang, Xinming, et al.
Veröffentlicht: (2026)
DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents
von: Zhang, Junshuo, et al.
Veröffentlicht: (2026)
von: Zhang, Junshuo, et al.
Veröffentlicht: (2026)
NCV: A Node-Wise Consistency Verification Approach for Low-Cost Structured Error Localization in LLM Reasoning
von: Zhang, Yulong, et al.
Veröffentlicht: (2025)
von: Zhang, Yulong, et al.
Veröffentlicht: (2025)
Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning
von: Xu, Ningning, et al.
Veröffentlicht: (2025)
von: Xu, Ningning, et al.
Veröffentlicht: (2025)
ToolDreamer: Instilling LLM Reasoning Into Tool Retrievers
von: Sengupta, Saptarshi, et al.
Veröffentlicht: (2025)
von: Sengupta, Saptarshi, et al.
Veröffentlicht: (2025)
Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning
von: Zhai, Naixin, et al.
Veröffentlicht: (2026)
von: Zhai, Naixin, et al.
Veröffentlicht: (2026)
Learning Wisdom from Errors: Promoting LLM's Continual Relation Learning through Exploiting Error Cases
von: Yin, Shaozhe, et al.
Veröffentlicht: (2025)
von: Yin, Shaozhe, et al.
Veröffentlicht: (2025)
Contextual Drag: How Errors in the Context Affect LLM Reasoning
von: Cheng, Yun, et al.
Veröffentlicht: (2026)
von: Cheng, Yun, et al.
Veröffentlicht: (2026)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
von: Lu, Meng, et al.
Veröffentlicht: (2025)
von: Lu, Meng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated Penalty
von: Yu, Zewei, et al.
Veröffentlicht: (2026) -
PORTool: Importance-Aware Policy Optimization with Rewarded Tree for Multi-Tool-Integrated Reasoning
von: Wu, Feijie, et al.
Veröffentlicht: (2025) -
AdaTIR: Adaptive Tool-Integrated Reasoning via Difficulty-Aware Policy Optimization
von: Fang, Zhaiyu, et al.
Veröffentlicht: (2026) -
Multi-Agent Tool-Integrated Policy Optimization
von: Mo, Zhanfeng, et al.
Veröffentlicht: (2025) -
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
von: Xu, Ran, et al.
Veröffentlicht: (2025)