Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Ruocheng, Dong, Kaiwen, Gao, Xiang, Das, Kamalika |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning
von: Gao, Xiang, et al.
Veröffentlicht: (2025)
von: Gao, Xiang, et al.
Veröffentlicht: (2025)
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
von: Li, Dawei, et al.
Veröffentlicht: (2026)
von: Li, Dawei, et al.
Veröffentlicht: (2026)
Customizing Language Model Responses with Contrastive In-Context Learning
von: Gao, Xiang, et al.
Veröffentlicht: (2024)
von: Gao, Xiang, et al.
Veröffentlicht: (2024)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
von: Baidya, Avinash, et al.
Veröffentlicht: (2025)
von: Baidya, Avinash, et al.
Veröffentlicht: (2025)
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
von: Lu, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Lu, Jiaxuan, et al.
Veröffentlicht: (2026)
Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents
von: Zhang, Kaituo, et al.
Veröffentlicht: (2026)
von: Zhang, Kaituo, et al.
Veröffentlicht: (2026)
UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents
von: Liang, Yijuan, et al.
Veröffentlicht: (2026)
von: Liang, Yijuan, et al.
Veröffentlicht: (2026)
Utility-Guided Agent Orchestration for Efficient LLM Tool Use
von: Liu, Boyan, et al.
Veröffentlicht: (2026)
von: Liu, Boyan, et al.
Veröffentlicht: (2026)
ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding
von: Liu, Xiao, et al.
Veröffentlicht: (2026)
von: Liu, Xiao, et al.
Veröffentlicht: (2026)
MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
von: Li, Binxu, et al.
Veröffentlicht: (2024)
von: Li, Binxu, et al.
Veröffentlicht: (2024)
Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution
von: He, Kaiwen, et al.
Veröffentlicht: (2025)
von: He, Kaiwen, et al.
Veröffentlicht: (2025)
Goal-Conditioned Supervised Learning for LLM Fine-Tuning
von: Li, Shijun, et al.
Veröffentlicht: (2026)
von: Li, Shijun, et al.
Veröffentlicht: (2026)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
von: Fei, Xiang, et al.
Veröffentlicht: (2025)
von: Fei, Xiang, et al.
Veröffentlicht: (2025)
When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems
von: Huang, Donghao, et al.
Veröffentlicht: (2026)
von: Huang, Donghao, et al.
Veröffentlicht: (2026)
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
von: Thaman, Kunvar
Veröffentlicht: (2026)
von: Thaman, Kunvar
Veröffentlicht: (2026)
ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use
von: Deng, Mengjie, et al.
Veröffentlicht: (2025)
von: Deng, Mengjie, et al.
Veröffentlicht: (2025)
TADDLE: A Tool-Augmented Agent for Detecting Deficient LLM-Generated Peer Reviews
von: Duan, Hanqi, et al.
Veröffentlicht: (2026)
von: Duan, Hanqi, et al.
Veröffentlicht: (2026)
Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use
von: Cheng, Yize, et al.
Veröffentlicht: (2026)
von: Cheng, Yize, et al.
Veröffentlicht: (2026)
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
von: Hu, Xuhao, et al.
Veröffentlicht: (2026)
von: Hu, Xuhao, et al.
Veröffentlicht: (2026)
When Does Memory Help Multi-Trajectory Inference for Tool-Use LLM Agents?
von: Li, Xinzhe, et al.
Veröffentlicht: (2026)
von: Li, Xinzhe, et al.
Veröffentlicht: (2026)
Factored Agents: Decoupling In-Context Learning and Memorization for Robust Tool Use
von: Roth, Nicholas, et al.
Veröffentlicht: (2025)
von: Roth, Nicholas, et al.
Veröffentlicht: (2025)
ToolTweak: An Attack on Tool Selection in LLM-based Agents
von: Sneh, Jonathan, et al.
Veröffentlicht: (2025)
von: Sneh, Jonathan, et al.
Veröffentlicht: (2025)
EvoTool: Self-Evolving Tool-Use Policy Optimization in LLM Agents via Blame-Aware Mutation and Diversity-Aware Selection
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
SPUQ: Perturbation-Based Uncertainty Quantification for Large Language Models
von: Gao, Xiang, et al.
Veröffentlicht: (2024)
von: Gao, Xiang, et al.
Veröffentlicht: (2024)
Guided by Trajectories: Repairing and Rewarding Tool-Use Trajectories for Tool-Integrated Reasoning
von: Gong, Siyu, et al.
Veröffentlicht: (2026)
von: Gong, Siyu, et al.
Veröffentlicht: (2026)
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
von: Feng, Jiazhan, et al.
Veröffentlicht: (2025)
von: Feng, Jiazhan, et al.
Veröffentlicht: (2025)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
von: Guo, Zhengkang, et al.
Veröffentlicht: (2026)
von: Guo, Zhengkang, et al.
Veröffentlicht: (2026)
Budget-Aware Tool-Use Enables Effective Agent Scaling
von: Liu, Tengxiao, et al.
Veröffentlicht: (2025)
von: Liu, Tengxiao, et al.
Veröffentlicht: (2025)
Tool-Augmented Policy Optimization: Synergizing Reasoning and Adaptive Tool Use with Reinforcement Learning
von: Wu, Wenxun, et al.
Veröffentlicht: (2025)
von: Wu, Wenxun, et al.
Veröffentlicht: (2025)
Reducing Cognitive Overhead in Tool Use via Multi-Small-Agent Reinforcement Learning
von: Wang, Dayu, et al.
Veröffentlicht: (2025)
von: Wang, Dayu, et al.
Veröffentlicht: (2025)
Benchmarking LLM Tool-Use in the Wild
von: Yu, Peijie, et al.
Veröffentlicht: (2026)
von: Yu, Peijie, et al.
Veröffentlicht: (2026)
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
von: Wang, Ruipeng, et al.
Veröffentlicht: (2026)
von: Wang, Ruipeng, et al.
Veröffentlicht: (2026)
CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval
von: Senthil, Vaishali, et al.
Veröffentlicht: (2026)
von: Senthil, Vaishali, et al.
Veröffentlicht: (2026)
LLM Agents Making Agent Tools
von: Wölflein, Georg, et al.
Veröffentlicht: (2025)
von: Wölflein, Georg, et al.
Veröffentlicht: (2025)
Node-Level Uncertainty Estimation in LLM-Generated SQL
von: Hasson, Hilaf, et al.
Veröffentlicht: (2025)
von: Hasson, Hilaf, et al.
Veröffentlicht: (2025)
PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools
von: Feng, Tianjun, et al.
Veröffentlicht: (2026)
von: Feng, Tianjun, et al.
Veröffentlicht: (2026)
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
GAP: Graph-Based Agent Planning with Parallel Tool Use and Reinforcement Learning
von: Wu, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wu, Jiaqi, et al.
Veröffentlicht: (2025)
Proof-of-Use: Mitigating Tool-Call Hacking in Deep Research Agents
von: Ma, SHengjie, et al.
Veröffentlicht: (2025)
von: Ma, SHengjie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning
von: Gao, Xiang, et al.
Veröffentlicht: (2025) -
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
von: Li, Dawei, et al.
Veröffentlicht: (2026) -
Customizing Language Model Responses with Contrastive In-Context Learning
von: Gao, Xiang, et al.
Veröffentlicht: (2024) -
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
von: Baidya, Avinash, et al.
Veröffentlicht: (2025) -
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
von: Lu, Jiaxuan, et al.
Veröffentlicht: (2026)