ToolMisuseBench: An Offline Deterministic Benchmark for Tool Misuse and Recovery in Agentic Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Sigdel, Akshey, Baral, Rista |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Schema First Tool APIs for LLM Agents: A Controlled Study of Tool Misuse, Recovery, and Budgeted Performance
by: Sigdel, Akshey, et al.
Published: (2026)
by: Sigdel, Akshey, et al.
Published: (2026)
Guardrails as Infrastructure: Policy-First Control for Tool-Orchestrated Workflows
by: Sigdel, Akshey, et al.
Published: (2026)
by: Sigdel, Akshey, et al.
Published: (2026)
EvolveTool-Bench: Evaluating the Quality of LLM-Generated Tool Libraries as Software Artifacts
by: Kaliyev, Alibek T., et al.
Published: (2026)
by: Kaliyev, Alibek T., et al.
Published: (2026)
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
by: Kokane, Shirley, et al.
Published: (2024)
by: Kokane, Shirley, et al.
Published: (2024)
On Generalization in Agentic Tool Calling: CoreThink Agentic Reasoner and MAVEN Dataset
by: Bhat, Vishvesh, et al.
Published: (2025)
by: Bhat, Vishvesh, et al.
Published: (2025)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
by: Zhou, Qixing, et al.
Published: (2026)
by: Zhou, Qixing, et al.
Published: (2026)
Applying an Agentic Coding Tool for Improving Published Algorithm Implementations
by: Suwannik, Worasait
Published: (2026)
by: Suwannik, Worasait
Published: (2026)
DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use
by: Chen, Aili, et al.
Published: (2026)
by: Chen, Aili, et al.
Published: (2026)
GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
by: Lindenbauer, Tobias, et al.
Published: (2025)
by: Lindenbauer, Tobias, et al.
Published: (2025)
Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
by: Elder, Benjamin, et al.
Published: (2025)
by: Elder, Benjamin, et al.
Published: (2025)
ToolFuzz -- Automated Agent Tool Testing
by: Milev, Ivan, et al.
Published: (2025)
by: Milev, Ivan, et al.
Published: (2025)
Investigating Tool-Memory Conflicts in Tool-Augmented LLMs
by: Cheng, Jiali, et al.
Published: (2026)
by: Cheng, Jiali, et al.
Published: (2026)
ParaTool: Shifting Tool Representations from Context to Parameters
by: Yu, Zekai, et al.
Published: (2026)
by: Yu, Zekai, et al.
Published: (2026)
TSCG: Deterministic Tool-Schema Compilation for Agentic LLM Deployments
by: Sakizli, Furkan
Published: (2026)
by: Sakizli, Furkan
Published: (2026)
Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Unified Approach for Elevating Benchmark Quality
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis
by: Qin, Bowen
Published: (2026)
by: Qin, Bowen
Published: (2026)
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
by: Li, Dawei, et al.
Published: (2026)
by: Li, Dawei, et al.
Published: (2026)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
by: Yang, Jie, et al.
Published: (2026)
by: Yang, Jie, et al.
Published: (2026)
An Executable Benchmarking Suite for Tool-Using Agents
by: Zhong, Zhiqing, et al.
Published: (2026)
by: Zhong, Zhiqing, et al.
Published: (2026)
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
by: Bandi, Chaithanya, et al.
Published: (2026)
by: Bandi, Chaithanya, et al.
Published: (2026)
Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems
by: Ouyang, Yipeng, et al.
Published: (2026)
by: Ouyang, Yipeng, et al.
Published: (2026)
Identifying and Mitigating API Misuse in Large Language Models
by: Zhuo, Terry Yue, et al.
Published: (2025)
by: Zhuo, Terry Yue, et al.
Published: (2025)
An Empirical Study of API Misuses of Data-Centric Libraries
by: Galappaththi, Akalanka, et al.
Published: (2024)
by: Galappaththi, Akalanka, et al.
Published: (2024)
REGAL: A Registry-Driven Architecture for Deterministic Grounding of Agentic AI in Enterprise Telemetry
by: Agrawal, Yuvraj
Published: (2026)
by: Agrawal, Yuvraj
Published: (2026)
The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge?
by: Zeng, Yirong, et al.
Published: (2026)
by: Zeng, Yirong, et al.
Published: (2026)
Semantic Tool Discovery for Large Language Models: A Vector-Based Approach to MCP Tool Selection
by: Mudunuri, Sarat, et al.
Published: (2026)
by: Mudunuri, Sarat, et al.
Published: (2026)
Breaking the Illusion of Identity in LLM Tooling
by: Miller, Marek
Published: (2026)
by: Miller, Marek
Published: (2026)
How the Misuse of a Dataset Harmed Semantic Clone Detection
by: Krinke, Jens, et al.
Published: (2025)
by: Krinke, Jens, et al.
Published: (2025)
MTAD: Tools and Benchmarks for Multivariate Time Series Anomaly Detection
by: Liu, Jinyang, et al.
Published: (2024)
by: Liu, Jinyang, et al.
Published: (2024)
User Centric Evaluation of Code Generation Tools
by: Miah, Tanha, et al.
Published: (2024)
by: Miah, Tanha, et al.
Published: (2024)
RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades
by: Xu, Xinbo, et al.
Published: (2026)
by: Xu, Xinbo, et al.
Published: (2026)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
by: Zhang, Wentao, et al.
Published: (2026)
by: Zhang, Wentao, et al.
Published: (2026)
VQA support to Arabic Language Learning Educational Tool
by: Delassi, Khaled Bachir, et al.
Published: (2025)
by: Delassi, Khaled Bachir, et al.
Published: (2025)
Tool-integrated Reinforcement Learning for Repo Deep Search
by: Ma, Zexiong, et al.
Published: (2025)
by: Ma, Zexiong, et al.
Published: (2025)
Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems
by: Xiong, Qian, et al.
Published: (2025)
by: Xiong, Qian, et al.
Published: (2025)
SWE Context Bench: A Benchmark for Context Learning in Coding
by: Zhu, Jiayuan, et al.
Published: (2026)
by: Zhu, Jiayuan, et al.
Published: (2026)
PBT-Bench: Benchmarking AI Agents on Property-Based Testing
by: Jing, Lucas, et al.
Published: (2026)
by: Jing, Lucas, et al.
Published: (2026)
AFGNN: API Misuse Detection using Graph Neural Networks and Clustering
by: Pirapuraj, Ponnampalam, et al.
Published: (2026)
by: Pirapuraj, Ponnampalam, et al.
Published: (2026)
Satellite: Detecting and Analyzing Smart Contract Vulnerabilities caused by Subcontract Misuse
by: Liao, Zeqin, et al.
Published: (2025)
by: Liao, Zeqin, et al.
Published: (2025)
ASA: Training-Free Representation Engineering for Tool-Calling Agents
by: Wang, Youjin, et al.
Published: (2026)
by: Wang, Youjin, et al.
Published: (2026)
Similar Items
-
Schema First Tool APIs for LLM Agents: A Controlled Study of Tool Misuse, Recovery, and Budgeted Performance
by: Sigdel, Akshey, et al.
Published: (2026) -
Guardrails as Infrastructure: Policy-First Control for Tool-Orchestrated Workflows
by: Sigdel, Akshey, et al.
Published: (2026) -
EvolveTool-Bench: Evaluating the Quality of LLM-Generated Tool Libraries as Software Artifacts
by: Kaliyev, Alibek T., et al.
Published: (2026) -
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
by: Kokane, Shirley, et al.
Published: (2024) -
On Generalization in Agentic Tool Calling: CoreThink Agentic Reasoner and MAVEN Dataset
by: Bhat, Vishvesh, et al.
Published: (2025)