ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Jiarui, Holleis, Thomas, Zhang, Yizhe, Aumayer, Bernhard, Nan, Feng, Bai, Felix, Ma, Shuang, Ma, Shen, Li, Mengyu, Yin, Guoli, Wang, Zirui, Pang, Ruoming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
von: Findeis, Arduin, et al.
Veröffentlicht: (2025)
von: Findeis, Arduin, et al.
Veröffentlicht: (2025)
MMAU: A Holistic Benchmark of Agent Capabilities Across Diverse Domains
von: Yin, Guoli, et al.
Veröffentlicht: (2024)
von: Yin, Guoli, et al.
Veröffentlicht: (2024)
Benchmarking LLM Tool-Use in the Wild
von: Yu, Peijie, et al.
Veröffentlicht: (2026)
von: Yu, Peijie, et al.
Veröffentlicht: (2026)
MCP-RADAR: A Multi-Dimensional Benchmark for Evaluating Tool Use Capabilities in Large Language Models
von: Gao, Xuanqi, et al.
Veröffentlicht: (2025)
von: Gao, Xuanqi, et al.
Veröffentlicht: (2025)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
von: Li, Yuanyang, et al.
Veröffentlicht: (2026)
von: Li, Yuanyang, et al.
Veröffentlicht: (2026)
RefTool: Reference-Guided Tool Creation for Knowledge-Intensive Reasoning
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
Quantifying Frontier LLM Capabilities for Container Sandbox Escape
von: Marchand, Rahul, et al.
Veröffentlicht: (2026)
von: Marchand, Rahul, et al.
Veröffentlicht: (2026)
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
von: Xu, Haoyuan, et al.
Veröffentlicht: (2026)
von: Xu, Haoyuan, et al.
Veröffentlicht: (2026)
Revisiting MoE and Dense Speed-Accuracy Comparisons for LLM Training
von: Du, Xianzhi, et al.
Veröffentlicht: (2024)
von: Du, Xianzhi, et al.
Veröffentlicht: (2024)
Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games
von: Zhang, Yizhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yizhe, et al.
Veröffentlicht: (2023)
Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning
von: Ma, Qinghe, et al.
Veröffentlicht: (2026)
von: Ma, Qinghe, et al.
Veröffentlicht: (2026)
Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities
von: Ying, Shuangshuang, et al.
Veröffentlicht: (2026)
von: Ying, Shuangshuang, et al.
Veröffentlicht: (2026)
TalkPlay-Tools: Conversational Music Recommendation with LLM Tool Calling
von: Doh, Seungheon, et al.
Veröffentlicht: (2025)
von: Doh, Seungheon, et al.
Veröffentlicht: (2025)
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
von: Thaman, Kunvar
Veröffentlicht: (2026)
von: Thaman, Kunvar
Veröffentlicht: (2026)
Instruction-Following Pruning for Large Language Models
von: Hou, Bairu, et al.
Veröffentlicht: (2025)
von: Hou, Bairu, et al.
Veröffentlicht: (2025)
Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use
von: Pang, Renning, et al.
Veröffentlicht: (2026)
von: Pang, Renning, et al.
Veröffentlicht: (2026)
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
von: Kokane, Shirley, et al.
Veröffentlicht: (2024)
von: Kokane, Shirley, et al.
Veröffentlicht: (2024)
AutoTool: Automatic Scaling of Tool-Use Capabilities in RL via Decoupled Entropy Constraints
von: Zeng, Yirong, et al.
Veröffentlicht: (2026)
von: Zeng, Yirong, et al.
Veröffentlicht: (2026)
Learning to Use Tools via Cooperative and Interactive Agents
von: Shi, Zhengliang, et al.
Veröffentlicht: (2024)
von: Shi, Zhengliang, et al.
Veröffentlicht: (2024)
SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
von: Jiang, Yanna, et al.
Veröffentlicht: (2026)
von: Jiang, Yanna, et al.
Veröffentlicht: (2026)
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
von: Huang, Yue, et al.
Veröffentlicht: (2023)
von: Huang, Yue, et al.
Veröffentlicht: (2023)
Die vergebliche Gabe
von: Holleis, Hans
Veröffentlicht: (2024)
von: Holleis, Hans
Veröffentlicht: (2024)
Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use
von: Guo, Ruocheng, et al.
Veröffentlicht: (2026)
von: Guo, Ruocheng, et al.
Veröffentlicht: (2026)
From Conversation to Query Execution: Benchmarking User and Tool Interactions for EHR Database Agents
von: Lee, Gyubok, et al.
Veröffentlicht: (2025)
von: Lee, Gyubok, et al.
Veröffentlicht: (2025)
Advancing SLM Tool-Use Capability using Reinforcement Learning
von: Paprunia, Dhruvi, et al.
Veröffentlicht: (2025)
von: Paprunia, Dhruvi, et al.
Veröffentlicht: (2025)
Controlling Performance and Budget of a Centralized Multi-agent LLM System with Reinforcement Learning
von: Jin, Bowen, et al.
Veröffentlicht: (2025)
von: Jin, Bowen, et al.
Veröffentlicht: (2025)
The Tool Illusion: Rethinking Tool Use in Web Agents
von: Lou, Renze, et al.
Veröffentlicht: (2026)
von: Lou, Renze, et al.
Veröffentlicht: (2026)
AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture
von: Ye, Zi, et al.
Veröffentlicht: (2026)
von: Ye, Zi, et al.
Veröffentlicht: (2026)
Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges
von: Wang, Hongru, et al.
Veröffentlicht: (2025)
von: Wang, Hongru, et al.
Veröffentlicht: (2025)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
von: Ma, Zixian, et al.
Veröffentlicht: (2024)
von: Ma, Zixian, et al.
Veröffentlicht: (2024)
ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
von: Liu, Marianne Menglin, et al.
Veröffentlicht: (2025)
von: Liu, Marianne Menglin, et al.
Veröffentlicht: (2025)
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
von: Xia, Bowei, et al.
Veröffentlicht: (2026)
von: Xia, Bowei, et al.
Veröffentlicht: (2026)
ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context
von: Xiu, Zidi, et al.
Veröffentlicht: (2026)
von: Xiu, Zidi, et al.
Veröffentlicht: (2026)
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
von: Shen, Yujiong, et al.
Veröffentlicht: (2026)
von: Shen, Yujiong, et al.
Veröffentlicht: (2026)
ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints
von: Choi, Hyeonje, et al.
Veröffentlicht: (2026)
von: Choi, Hyeonje, et al.
Veröffentlicht: (2026)
Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools
von: Lymperopoulos, Panagiotis, et al.
Veröffentlicht: (2025)
von: Lymperopoulos, Panagiotis, et al.
Veröffentlicht: (2025)
Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents
von: Zhang, Kaituo, et al.
Veröffentlicht: (2026)
von: Zhang, Kaituo, et al.
Veröffentlicht: (2026)
Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo
von: Feng, Shengyu, et al.
Veröffentlicht: (2024)
von: Feng, Shengyu, et al.
Veröffentlicht: (2024)
EvoTool: Self-Evolving Tool-Use Policy Optimization in LLM Agents via Blame-Aware Mutation and Diversity-Aware Selection
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
von: Findeis, Arduin, et al.
Veröffentlicht: (2025) -
MMAU: A Holistic Benchmark of Agent Capabilities Across Diverse Domains
von: Yin, Guoli, et al.
Veröffentlicht: (2024) -
Benchmarking LLM Tool-Use in the Wild
von: Yu, Peijie, et al.
Veröffentlicht: (2026) -
MCP-RADAR: A Multi-Dimensional Benchmark for Evaluating Tool Use Capabilities in Large Language Models
von: Gao, Xuanqi, et al.
Veröffentlicht: (2025) -
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
von: Li, Yuanyang, et al.
Veröffentlicht: (2026)