Learning from the Irrecoverable: Error-Localized Policy Optimization for Tool-Integrated LLM Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Qiao, Zhu, Yuke, Ge, Chao, Yang, Lei, Shen, Ying, Zheng, Bo, Guo, Sheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914319043657728
author Liang, Qiao
Zhu, Yuke
Ge, Chao
Yang, Lei
Shen, Ying
Zheng, Bo
Guo, Sheng
author_facet Liang, Qiao
Zhu, Yuke
Ge, Chao
Yang, Lei
Shen, Ying
Zheng, Bo
Guo, Sheng
contents Tool-integrated reasoning (TIR) enables LLM agents to solve tasks through planning, tool use, and iterative revision, but outcome-only reinforcement learning in this setting suffers from sparse, delayed rewards and weak step-level credit assignment. In long-horizon TIR trajectories, an early irrecoverable mistake can determine success or failure, making it crucial to localize the first irrecoverable step and leverage it for fine-grained credit assignment. We propose Error-Localized Policy Optimization (ELPO), which localizes the first irrecoverable step via binary-search rollout trees under a fixed rollout budget, converts the resulting tree into stable learning signals through hierarchical advantage attribution, and applies error-localized adaptive clipping to strengthen corrective updates on the critical step and its suffix. Across TIR benchmarks in math, science QA, and code execution, ELPO consistently outperforms strong Agentic RL baselines under comparable sampling budgets, with additional gains in Pass@K and Major@K scaling, rollout ranking quality, and tool-call efficiency. Our code will be publicly released soon.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09598
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning from the Irrecoverable: Error-Localized Policy Optimization for Tool-Integrated LLM Reasoning
Liang, Qiao
Zhu, Yuke
Ge, Chao
Yang, Lei
Shen, Ying
Zheng, Bo
Guo, Sheng
Computation and Language
Tool-integrated reasoning (TIR) enables LLM agents to solve tasks through planning, tool use, and iterative revision, but outcome-only reinforcement learning in this setting suffers from sparse, delayed rewards and weak step-level credit assignment. In long-horizon TIR trajectories, an early irrecoverable mistake can determine success or failure, making it crucial to localize the first irrecoverable step and leverage it for fine-grained credit assignment. We propose Error-Localized Policy Optimization (ELPO), which localizes the first irrecoverable step via binary-search rollout trees under a fixed rollout budget, converts the resulting tree into stable learning signals through hierarchical advantage attribution, and applies error-localized adaptive clipping to strengthen corrective updates on the critical step and its suffix. Across TIR benchmarks in math, science QA, and code execution, ELPO consistently outperforms strong Agentic RL baselines under comparable sampling budgets, with additional gains in Pass@K and Major@K scaling, rollout ranking quality, and tool-call efficiency. Our code will be publicly released soon.
title Learning from the Irrecoverable: Error-Localized Policy Optimization for Tool-Integrated LLM Reasoning
topic Computation and Language
url https://arxiv.org/abs/2602.09598