MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Wenrui, Liu, Zixiang, Dai, Elsie, Yu, Wenhan, Yu, Lei, Yang, Tong, Han, Jinjun, Gao, Hong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918297781403648
author Liu, Wenrui
Liu, Zixiang
Dai, Elsie
Yu, Wenhan
Yu, Lei
Yang, Tong
Han, Jinjun
Gao, Hong
author_facet Liu, Wenrui
Liu, Zixiang
Dai, Elsie
Yu, Wenhan
Yu, Lei
Yang, Tong
Han, Jinjun
Gao, Hong
contents Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a future trend. Current MCP evaluation sets suffer from issues such as reliance on external MCP services and a lack of difficulty awareness. To address these limitations, we propose MCPAgentBench, a benchmark based on real-world MCP definitions designed to evaluate the tool-use capabilities of agents. We construct a dataset containing authentic tasks and simulated MCP tools. The evaluation employs a dynamic sandbox environment that presents agents with candidate tool lists containing distractors, thereby testing their tool selection and discrimination abilities. Furthermore, we introduce comprehensive metrics to measure both task completion rates and execution efficiency. Experiments conducted on various latest mainstream Large Language Models reveal significant performance differences in handling complex, multi-step tool invocations. All code is open-source at Github.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24565
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
Liu, Wenrui
Liu, Zixiang
Dai, Elsie
Yu, Wenhan
Yu, Lei
Yang, Tong
Han, Jinjun
Gao, Hong
Artificial Intelligence
Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a future trend. Current MCP evaluation sets suffer from issues such as reliance on external MCP services and a lack of difficulty awareness. To address these limitations, we propose MCPAgentBench, a benchmark based on real-world MCP definitions designed to evaluate the tool-use capabilities of agents. We construct a dataset containing authentic tasks and simulated MCP tools. The evaluation employs a dynamic sandbox environment that presents agents with candidate tool lists containing distractors, thereby testing their tool selection and discrimination abilities. Furthermore, we introduce comprehensive metrics to measure both task completion rates and execution efficiency. Experiments conducted on various latest mainstream Large Language Models reveal significant performance differences in handling complex, multi-step tool invocations. All code is open-source at Github.
title MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
topic Artificial Intelligence
url https://arxiv.org/abs/2512.24565