JTPRO: A Joint Tool-Prompt Reflective Optimization Framework for Language Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghoshal, Sandip, Mittal, Anshul, Singh, Jyotika, Ballesteros, Miguel, Sun, Weiyi, Tu, Fang, Singh, Shailender, Benajiba, Yassine, Shah, Fahad, Bharadwaj, Sujeeth, Ravi, Sujith, Roth, Dan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913053504700416
author Ghoshal, Sandip
Mittal, Anshul
Singh, Jyotika
Ballesteros, Miguel
Sun, Weiyi
Tu, Fang
Singh, Shailender
Benajiba, Yassine
Shah, Fahad
Bharadwaj, Sujeeth
Ravi, Sujith
Roth, Dan
author_facet Ghoshal, Sandip
Mittal, Anshul
Singh, Jyotika
Ballesteros, Miguel
Sun, Weiyi
Tu, Fang
Singh, Shailender
Benajiba, Yassine
Shah, Fahad
Bharadwaj, Sujeeth
Ravi, Sujith
Roth, Dan
contents Large language model (LLM) agents augmented with external tools often struggle as number of tools grow large and become domain-specific. In such settings, ambiguous tool descriptions and under-specified agent instructions frequently lead to tool mis-selection and incorrect slot/value instantiation. We hypothesize that this is due to two root causes: generic, one-size-fits-all prompts that ignore tool-specific nuances, and underspecified tool schemas that lack clear guidance on when and how to use each tool and how to format its parameters. We introduce Joint Tool-Prompt Reflective Optimization (JTPRO), a framework for improving tool-calling reliability in trace-supervised settings by iteratively using rollout-driven reflection to co-optimize global instructions and per-tool schema/argument descriptions for accurate tool selection and argument instantiation in large tool inventories. JTPRO is designed to preserve only tool-local cues needed for correct disambiguation and slot filling. We evaluate JTPRO across multi-tool benchmarks, which account for different number of tools using three metrics: Tool Selection Accuracy (TSA), Slot Filling Accuracy(SFA), and Overall Success Rate(OSR) (correct tool + correct slots + correct values). JTPRO consistently outperforms strong baselines, including CoT-style agents, and reflective prompt optimizers such as GEPA by 5%-20% (relative) on OSR. Ablations show that joint optimization of instructions and tool schemas is more effective and robust than optimizing either component in isolation.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19821
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle JTPRO: A Joint Tool-Prompt Reflective Optimization Framework for Language Agents
Ghoshal, Sandip
Mittal, Anshul
Singh, Jyotika
Ballesteros, Miguel
Sun, Weiyi
Tu, Fang
Singh, Shailender
Benajiba, Yassine
Shah, Fahad
Bharadwaj, Sujeeth
Ravi, Sujith
Roth, Dan
Artificial Intelligence
Software Engineering
Large language model (LLM) agents augmented with external tools often struggle as number of tools grow large and become domain-specific. In such settings, ambiguous tool descriptions and under-specified agent instructions frequently lead to tool mis-selection and incorrect slot/value instantiation. We hypothesize that this is due to two root causes: generic, one-size-fits-all prompts that ignore tool-specific nuances, and underspecified tool schemas that lack clear guidance on when and how to use each tool and how to format its parameters. We introduce Joint Tool-Prompt Reflective Optimization (JTPRO), a framework for improving tool-calling reliability in trace-supervised settings by iteratively using rollout-driven reflection to co-optimize global instructions and per-tool schema/argument descriptions for accurate tool selection and argument instantiation in large tool inventories. JTPRO is designed to preserve only tool-local cues needed for correct disambiguation and slot filling. We evaluate JTPRO across multi-tool benchmarks, which account for different number of tools using three metrics: Tool Selection Accuracy (TSA), Slot Filling Accuracy(SFA), and Overall Success Rate(OSR) (correct tool + correct slots + correct values). JTPRO consistently outperforms strong baselines, including CoT-style agents, and reflective prompt optimizers such as GEPA by 5%-20% (relative) on OSR. Ablations show that joint optimization of instructions and tool schemas is more effective and robust than optimizing either component in isolation.
title JTPRO: A Joint Tool-Prompt Reflective Optimization Framework for Language Agents
topic Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2604.19821