Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhiwei, Zhao, Fei, Wang, Rui, Wang, Zezhong, Liang, Bin, Wang, Jiakang, Hu, Yao, Cao, Shaosheng, Wong, Kam-Fai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915944196997120
author Zhang, Zhiwei
Zhao, Fei
Wang, Rui
Wang, Zezhong
Liang, Bin
Wang, Jiakang
Hu, Yao
Cao, Shaosheng
Wong, Kam-Fai
author_facet Zhang, Zhiwei
Zhao, Fei
Wang, Rui
Wang, Zezhong
Liang, Bin
Wang, Jiakang
Hu, Yao
Cao, Shaosheng
Wong, Kam-Fai
contents Large language models (LLMs) can call tools effectively, yet they remain brittle in multi-turn execution: after a tool-call error, smaller models often fall into repetitive invalid re-invocations instead of interpreting the feedback and recovering. This failure mode persists because current training paradigms do not explicitly teach models how to recover from execution errors. In particular, standard reinforcement learning (RL) collapses rich failure experience into sparse negative rewards, while pre-collected error-correction datasets become mismatched to the policy's evolving failure modes. To bridge this gap, we propose Fission-GRPO, a framework that converts execution errors into on-policy corrective supervision within the RL training loop. Our core mechanism fissions each failed trajectory into a new training instance by augmenting it with diagnostic feedback from a fine-tuned Error Simulator, then resampling multiple recovery rollouts on-policy. This enables the model to learn from the precise errors it makes during exploration, rather than from static, pre-collected error cases. On BFCL v4 Multi-Turn, Fission-GRPO improves the error recovery rate of Qwen3-8B by 5.7% absolute and overall accuracy by 4.0% (from 42.75% to 46.75%), outperforming both RL baselines and specialized tool-use agents. The method further generalizes to TAU-Bench and TAU2-Bench, achieving leading results across most settings with gains up to +17.4%.
format Preprint
id arxiv_https___arxiv_org_abs_2601_15625
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors
Zhang, Zhiwei
Zhao, Fei
Wang, Rui
Wang, Zezhong
Liang, Bin
Wang, Jiakang
Hu, Yao
Cao, Shaosheng
Wong, Kam-Fai
Machine Learning
Artificial Intelligence
I.2.7
Large language models (LLMs) can call tools effectively, yet they remain brittle in multi-turn execution: after a tool-call error, smaller models often fall into repetitive invalid re-invocations instead of interpreting the feedback and recovering. This failure mode persists because current training paradigms do not explicitly teach models how to recover from execution errors. In particular, standard reinforcement learning (RL) collapses rich failure experience into sparse negative rewards, while pre-collected error-correction datasets become mismatched to the policy's evolving failure modes. To bridge this gap, we propose Fission-GRPO, a framework that converts execution errors into on-policy corrective supervision within the RL training loop. Our core mechanism fissions each failed trajectory into a new training instance by augmenting it with diagnostic feedback from a fine-tuned Error Simulator, then resampling multiple recovery rollouts on-policy. This enables the model to learn from the precise errors it makes during exploration, rather than from static, pre-collected error cases. On BFCL v4 Multi-Turn, Fission-GRPO improves the error recovery rate of Qwen3-8B by 5.7% absolute and overall accuracy by 4.0% (from 42.75% to 46.75%), outperforming both RL baselines and specialized tool-use agents. The method further generalizes to TAU-Bench and TAU2-Bench, achieving leading results across most settings with gains up to +17.4%.
title Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors
topic Machine Learning
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2601.15625