PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Luan, Song, Dandan, Wu, Zhijing, Chen, Zhengyu, Zhang, Chen, Tian, Yuhang, Ma, Huipeng, Li, Chenhao, Zhou, Changzhi, Li, Xudong, Zhang, Shuhao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913110403579904
author Zhang, Luan
Song, Dandan
Wu, Zhijing
Chen, Zhengyu
Zhang, Chen
Tian, Yuhang
Ma, Huipeng
Li, Chenhao
Zhou, Changzhi
Li, Xudong
Zhang, Shuhao
author_facet Zhang, Luan
Song, Dandan
Wu, Zhijing
Chen, Zhengyu
Zhang, Chen
Tian, Yuhang
Ma, Huipeng
Li, Chenhao
Zhou, Changzhi
Li, Xudong
Zhang, Shuhao
contents Tool-integrated reasoning (TIR) enables large language models (LLMs) to enhance their capabilities by interacting with external tools, such as code interpreters (CI). Most recent studies focus on exploring various methods to equip LLMs with the ability to use tools. However, how to further boost the reasoning ability of already tool-capable LLMs at inference time remains underexplored. Improving reasoning at inference time requires no additional training and can help LLMs better leverage tools to solve problems. We observe that, during tool-capable LLM inference, both the number and the proportion of erroneous tool calls are negatively correlated with answer correctness. Moreover, erroneous tool calls are typically resolved successfully within a few subsequent turns. If not, LLMs often struggle to resolve such errors even with many additional turns. Building on the above observations, we propose PruneTIR, a rather effective yet efficient framework that enhances the tool-integrated reasoning at inference time. During LLM inference, PruneTIR prunes trajectories, resamples tool calls, and suspends tool usage through three components: Success-Triggered Pruning, Stuck-Triggered Pruning and Resampling, and Retry-Triggered Tool Suspension. These three components enable PruneTIR to mitigate the negative impact of erroneous tool calls and prevent LLMs from getting stuck in repeated failed resolution attempts, thereby improving overall LLM performance. Extensive experimental results demonstrate the effectiveness of PruneTIR, which significantly improves Pass@1 and efficiency while reducing the working context length for tool-capable LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_09931
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning
Zhang, Luan
Song, Dandan
Wu, Zhijing
Chen, Zhengyu
Zhang, Chen
Tian, Yuhang
Ma, Huipeng
Li, Chenhao
Zhou, Changzhi
Li, Xudong
Zhang, Shuhao
Computation and Language
Artificial Intelligence
Tool-integrated reasoning (TIR) enables large language models (LLMs) to enhance their capabilities by interacting with external tools, such as code interpreters (CI). Most recent studies focus on exploring various methods to equip LLMs with the ability to use tools. However, how to further boost the reasoning ability of already tool-capable LLMs at inference time remains underexplored. Improving reasoning at inference time requires no additional training and can help LLMs better leverage tools to solve problems. We observe that, during tool-capable LLM inference, both the number and the proportion of erroneous tool calls are negatively correlated with answer correctness. Moreover, erroneous tool calls are typically resolved successfully within a few subsequent turns. If not, LLMs often struggle to resolve such errors even with many additional turns. Building on the above observations, we propose PruneTIR, a rather effective yet efficient framework that enhances the tool-integrated reasoning at inference time. During LLM inference, PruneTIR prunes trajectories, resamples tool calls, and suspends tool usage through three components: Success-Triggered Pruning, Stuck-Triggered Pruning and Resampling, and Retry-Triggered Tool Suspension. These three components enable PruneTIR to mitigate the negative impact of erroneous tool calls and prevent LLMs from getting stuck in repeated failed resolution attempts, thereby improving overall LLM performance. Extensive experimental results demonstrate the effectiveness of PruneTIR, which significantly improves Pass@1 and efficiency while reducing the working context length for tool-capable LLMs.
title PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.09931