MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhenting, Chang, Qi, Patel, Hemani, Biju, Shashank, Wu, Cheng-En, Liu, Quan, Ding, Aolin, Rezazadeh, Alireza, Shah, Ankit, Bao, Yujia, Siow, Eugene |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
by: Zong, Xuanjun, et al.
Published: (2025)
by: Zong, Xuanjun, et al.
Published: (2025)
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
by: Bandi, Chaithanya, et al.
Published: (2026)
by: Bandi, Chaithanya, et al.
Published: (2026)
MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
by: Wang, Zhiqiang, et al.
Published: (2025)
by: Wang, Zhiqiang, et al.
Published: (2025)
MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments
by: Ganapavarapu, Giridhar, et al.
Published: (2026)
by: Ganapavarapu, Giridhar, et al.
Published: (2026)
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
by: Guo, Zikang, et al.
Published: (2025)
by: Guo, Zikang, et al.
Published: (2025)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
by: Liu, Wenrui, et al.
Published: (2025)
by: Liu, Wenrui, et al.
Published: (2025)
Auditing MCP Servers for Over-Privileged Tool Capabilities
by: Huang, Charoes, et al.
Published: (2026)
by: Huang, Charoes, et al.
Published: (2026)
MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
by: Luo, Ziyang, et al.
Published: (2025)
by: Luo, Ziyang, et al.
Published: (2025)
Evaluation Report on MCP Servers
by: Luo, Zhiling, et al.
Published: (2025)
by: Luo, Zhiling, et al.
Published: (2025)
MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP Tools
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
by: Jia, Hongrui, et al.
Published: (2025)
by: Jia, Hongrui, et al.
Published: (2025)
MCP Pitfall Lab: Exposing Developer Pitfalls in MCP Tool Server Security under Multi-Vector Attacks
by: Hao, Run, et al.
Published: (2026)
by: Hao, Run, et al.
Published: (2026)
Scripts for measurement of MCP Server Marketplaces
by: Anonymous, Anonymous
Published: (2026)
by: Anonymous, Anonymous
Published: (2026)
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
A First Measurement Study on Authentication Security in Real-World Remote MCP Servers
by: Zhou, Huijun, et al.
Published: (2026)
by: Zhou, Huijun, et al.
Published: (2026)
TrialMCP Pack: MCP Servers for Physical AI Oncology Clinical Trial Systems
by: Kevin Kawchak, et al.
Published: (2026)
by: Kevin Kawchak, et al.
Published: (2026)
Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
MCPGuard : Automatically Detecting Vulnerabilities in MCP Servers
by: Wang, Bin, et al.
Published: (2025)
by: Wang, Bin, et al.
Published: (2025)
MCP-ITP: An Automated Framework for Implicit Tool Poisoning in MCP
by: Li, Ruiqi, et al.
Published: (2026)
by: Li, Ruiqi, et al.
Published: (2026)
When MCP Servers Attack: Taxonomy, Feasibility, and Mitigation
by: Zhao, Weibo, et al.
Published: (2025)
by: Zhao, Weibo, et al.
Published: (2025)
From Language to Action: Enhancing LLM Task Efficiency with Task-Aware MCP Server Recommendation
by: He, Shiyu, et al.
Published: (2026)
by: He, Shiyu, et al.
Published: (2026)
Trivial Trojans: How Minimal MCP Servers Enable Cross-Tool Exfiltration of Sensitive Data
by: Croce, Nicola, et al.
Published: (2025)
by: Croce, Nicola, et al.
Published: (2025)
Dynamic ReAct: Scalable Tool Selection for Large-Scale MCP Environments
by: Gaurav, Nishant, et al.
Published: (2025)
by: Gaurav, Nishant, et al.
Published: (2025)
MCP-SandboxScan: WASM-based Secure Execution and Runtime Analysis for MCP Tools
by: Tan, Zhuoran, et al.
Published: (2026)
by: Tan, Zhuoran, et al.
Published: (2026)
Enhancing Model Context Protocol (MCP) with Context-Aware Server Collaboration
by: Jayanti, Meenakshi Amulya, et al.
Published: (2026)
by: Jayanti, Meenakshi Amulya, et al.
Published: (2026)
From Docs to Descriptions: Smell-Aware Evaluation of MCP Server Descriptions
by: Wang, Peiran, et al.
Published: (2026)
by: Wang, Peiran, et al.
Published: (2026)
LiteCUA: Computer as MCP Server for Computer-Use Agent on AIOS
by: Mei, Kai, et al.
Published: (2025)
by: Mei, Kai, et al.
Published: (2025)
citecheck: An MCP Server for Automated Bibliographic Verification and Repair in Scholarly Manuscripts
by: Lee, Junhyeok
Published: (2026)
by: Lee, Junhyeok
Published: (2026)
MCPToolBench++: A Large Scale AI Agent Model Context Protocol MCP Tool Use Benchmark
by: Fan, Shiqing, et al.
Published: (2025)
by: Fan, Shiqing, et al.
Published: (2025)
HumanMCP: A Human-Like Query Dataset for Evaluating MCP Tool Retrieval Performance
by: Laddha, Shubh, et al.
Published: (2025)
by: Laddha, Shubh, et al.
Published: (2025)
MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation
by: Wang, Wenhao, et al.
Published: (2026)
by: Wang, Wenhao, et al.
Published: (2026)
TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration
by: Zhu, Yakun, et al.
Published: (2026)
by: Zhu, Yakun, et al.
Published: (2026)
A Modular Reference Architecture for MCP-Servers Enabling Agentic BIM Interaction
by: Heimig-Elschner, Tobias, et al.
Published: (2025)
by: Heimig-Elschner, Tobias, et al.
Published: (2025)
Can LLM Infer Risk Information From MCP Server System Logs?
by: Fu, Jiayi, et al.
Published: (2025)
by: Fu, Jiayi, et al.
Published: (2025)
From Component Manipulation to System Compromise: Understanding and Detecting Malicious MCP Servers
by: Huang, Yiheng, et al.
Published: (2026)
by: Huang, Yiheng, et al.
Published: (2026)
Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions
by: Hasan, Mohammed Mehedi, et al.
Published: (2026)
by: Hasan, Mohammed Mehedi, et al.
Published: (2026)
From Isolated Conversations to Hierarchical Schemas: Dynamic Tree Memory Representation for LLMs
by: Rezazadeh, Alireza, et al.
Published: (2024)
by: Rezazadeh, Alireza, et al.
Published: (2024)
Busemann and MCP
by: Fujioka, Tadashi, et al.
Published: (2026)
by: Fujioka, Tadashi, et al.
Published: (2026)
Code2MCP: Transforming Code Repositories into MCP Services
by: Ouyang, Chaoqian, et al.
Published: (2025)
by: Ouyang, Chaoqian, et al.
Published: (2025)
Similar Items
-
MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
by: Zong, Xuanjun, et al.
Published: (2025) -
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
by: Bandi, Chaithanya, et al.
Published: (2026) -
MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
by: Wang, Zhiqiang, et al.
Published: (2025) -
MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments
by: Ganapavarapu, Giridhar, et al.
Published: (2026) -
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
by: Guo, Zikang, et al.
Published: (2025)