MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yue, Shi, Jiawen, Li, Yuan, Fan, Chenrui, Wu, Siyuan, Zhang, Qihui, Liu, Yixin, Zhou, Pan, Wan, Yao, Gong, Neil Zhenqiang, Sun, Lichao |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
by: Kokane, Shirley, et al.
Published: (2024)
by: Kokane, Shirley, et al.
Published: (2024)
Prompt Injection Attack to Tool Selection in LLM Agents
by: Shi, Jiawen, et al.
Published: (2025)
by: Shi, Jiawen, et al.
Published: (2025)
Online-Optimized RAG for Tool Use and Function Calling
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints
by: Choi, Hyeonje, et al.
Published: (2026)
by: Choi, Hyeonje, et al.
Published: (2026)
Does Your Neural Code Completion Model Use My Code? A Membership Inference Approach
by: Wan, Yao, et al.
Published: (2024)
by: Wan, Yao, et al.
Published: (2024)
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
by: Xu, Haoyuan, et al.
Published: (2026)
by: Xu, Haoyuan, et al.
Published: (2026)
Towards Verifiably Safe Tool Use for LLM Agents
by: Doshi, Aarya, et al.
Published: (2026)
by: Doshi, Aarya, et al.
Published: (2026)
ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
by: Liu, Marianne Menglin, et al.
Published: (2025)
by: Liu, Marianne Menglin, et al.
Published: (2025)
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
by: Bandi, Chaithanya, et al.
Published: (2026)
by: Bandi, Chaithanya, et al.
Published: (2026)
Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
by: Li, Zeping, et al.
Published: (2026)
by: Li, Zeping, et al.
Published: (2026)
MetaTool: Facilitating Large Language Models to Master Tools with Meta-task Augmentation
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
AI Tool Use and Adoption in Software Development by Individuals and Organizations: A Grounded Theory Study
by: Li, Ze Shi, et al.
Published: (2024)
by: Li, Ze Shi, et al.
Published: (2024)
Prompting in Practice: Investigating Software Practitioners' Use of Generative AI Tools
by: Otten, Daniel, et al.
Published: (2025)
by: Otten, Daniel, et al.
Published: (2025)
Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation
by: Wang, Chong, et al.
Published: (2024)
by: Wang, Chong, et al.
Published: (2024)
Evaluating Tool Cloning in Agentic-AI Ecosystems
by: Kim, Taein, et al.
Published: (2026)
by: Kim, Taein, et al.
Published: (2026)
Self-Cognition in Large Language Models: An Exploratory Study
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
Trajectory Supervision for Continual Tool-Use Learning in LLMs
by: Reddy, Vishnu Vardhan, et al.
Published: (2026)
by: Reddy, Vishnu Vardhan, et al.
Published: (2026)
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
by: Yuan, Zenghui, et al.
Published: (2025)
by: Yuan, Zenghui, et al.
Published: (2025)
Can You Mimic Me? Exploring the Use of Android Record & Replay Tools in Debugging
by: Song, Zihe, et al.
Published: (2025)
by: Song, Zihe, et al.
Published: (2025)
Automated Tool Support for Category-Partition Testing: Design Decisions, UI and Examples of Use
by: Labiche, Yvan
Published: (2026)
by: Labiche, Yvan
Published: (2026)
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
by: Chen, Shiqi, et al.
Published: (2026)
by: Chen, Shiqi, et al.
Published: (2026)
DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use
by: Chen, Aili, et al.
Published: (2026)
by: Chen, Aili, et al.
Published: (2026)
Structure versus Context: Understanding the Design and Use of Computer Tools in Social Settings.
by: Powell, Kevin
Published: (1999)
by: Powell, Kevin
Published: (1999)
SMARTCAL: An Approach to Self-Aware Tool-Use Evaluation and Calibration
by: Shen, Yuanhao, et al.
Published: (2024)
by: Shen, Yuanhao, et al.
Published: (2024)
"I Don't Use AI for Everything": Exploring Utility, Attitude, and Responsibility of AI-empowered Tools in Software Development
by: Pan, Shidong, et al.
Published: (2024)
by: Pan, Shidong, et al.
Published: (2024)
Investigating Tool-Memory Conflicts in Tool-Augmented LLMs
by: Cheng, Jiali, et al.
Published: (2026)
by: Cheng, Jiali, et al.
Published: (2026)
"Maybe We Need Some More Examples:" Individual and Team Drivers of Developer GenAI Tool Use
by: Miller, Courtney, et al.
Published: (2025)
by: Miller, Courtney, et al.
Published: (2025)
Use as Directed? A Comparison of Software Tools Intended to Check Rigor and Transparency of Published Work
by: Eckmann, Peter, et al.
Published: (2025)
by: Eckmann, Peter, et al.
Published: (2025)
RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
by: LeVine, Will, et al.
Published: (2026)
by: LeVine, Will, et al.
Published: (2026)
A Soundness and Precision Benchmark for Java Debloating Tools
by: Klauke, Jonas, et al.
Published: (2025)
by: Klauke, Jonas, et al.
Published: (2025)
ToolMisuseBench: An Offline Deterministic Benchmark for Tool Misuse and Recovery in Agentic Systems
by: Sigdel, Akshey, et al.
Published: (2026)
by: Sigdel, Akshey, et al.
Published: (2026)
Which Prompting Technique Should I Use? An Empirical Investigation of Prompting Techniques for Software Engineering Tasks
by: Santana Jr, E. G., et al.
Published: (2025)
by: Santana Jr, E. G., et al.
Published: (2025)
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
by: Xia, Hongfei, et al.
Published: (2025)
by: Xia, Hongfei, et al.
Published: (2025)
Semantic Tool Discovery for Large Language Models: A Vector-Based Approach to MCP Tool Selection
by: Mudunuri, Sarat, et al.
Published: (2026)
by: Mudunuri, Sarat, et al.
Published: (2026)
Benchmarking Failures in Tool-Augmented Language Models
by: Treviño, Eduardo, et al.
Published: (2025)
by: Treviño, Eduardo, et al.
Published: (2025)
ParaTool: Shifting Tool Representations from Context to Parameters
by: Yu, Zekai, et al.
Published: (2026)
by: Yu, Zekai, et al.
Published: (2026)
A Novel Refactoring and Semantic Aware Abstract Syntax Tree Differencing Tool and a Benchmark for Evaluating the Accuracy of Diff Tools
by: Alikhanifard, Pouria, et al.
Published: (2024)
by: Alikhanifard, Pouria, et al.
Published: (2024)
UCRBench: Benchmarking LLMs on Use Case Recovery
by: Xiao, Shuyuan, et al.
Published: (2025)
by: Xiao, Shuyuan, et al.
Published: (2025)
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
by: Li, Dawei, et al.
Published: (2026)
by: Li, Dawei, et al.
Published: (2026)
ToolRosella: Translating Code Repositories into Standardized Tools for Scientific Agents
by: Di, Shimin, et al.
Published: (2026)
by: Di, Shimin, et al.
Published: (2026)
Similar Items
-
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
by: Kokane, Shirley, et al.
Published: (2024) -
Prompt Injection Attack to Tool Selection in LLM Agents
by: Shi, Jiawen, et al.
Published: (2025) -
Online-Optimized RAG for Tool Use and Function Calling
by: Pan, Yu, et al.
Published: (2025) -
ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints
by: Choi, Hyeonje, et al.
Published: (2026) -
Does Your Neural Code Completion Model Use My Code? A Membership Inference Approach
by: Wan, Yao, et al.
Published: (2024)