MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Wenrui, Liu, Zixiang, Dai, Elsie, Yu, Wenhan, Yu, Lei, Yang, Tong, Han, Jinjun, Gao, Hong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
by: Wang, Zhenting, et al.
Published: (2025)
by: Wang, Zhenting, et al.
Published: (2025)
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
by: Guo, Zikang, et al.
Published: (2025)
by: Guo, Zikang, et al.
Published: (2025)
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
by: Jia, Hongrui, et al.
Published: (2025)
by: Jia, Hongrui, et al.
Published: (2025)
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
by: He, Wei, et al.
Published: (2025)
by: He, Wei, et al.
Published: (2025)
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
by: Lu, Jiaxuan, et al.
Published: (2026)
by: Lu, Jiaxuan, et al.
Published: (2026)
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
by: Bandi, Chaithanya, et al.
Published: (2026)
by: Bandi, Chaithanya, et al.
Published: (2026)
MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
by: Zong, Xuanjun, et al.
Published: (2025)
by: Zong, Xuanjun, et al.
Published: (2025)
StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
by: Chen, Yanxu, et al.
Published: (2025)
by: Chen, Yanxu, et al.
Published: (2025)
MCPToolBench++: A Large Scale AI Agent Model Context Protocol MCP Tool Use Benchmark
by: Fan, Shiqing, et al.
Published: (2025)
by: Fan, Shiqing, et al.
Published: (2025)
EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents
by: Liu, Yunqi, et al.
Published: (2026)
by: Liu, Yunqi, et al.
Published: (2026)
MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP Tools
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
by: Zhang, Dongsen, et al.
Published: (2025)
by: Zhang, Dongsen, et al.
Published: (2025)
TPS-Bench: Evaluating AI Agents' Tool Planning \& Scheduling Abilities in Compounding Tasks
by: Xu, Hanwen, et al.
Published: (2025)
by: Xu, Hanwen, et al.
Published: (2025)
ParaView-MCP: An Autonomous Visualization Agent with Direct Tool Use
by: Liu, Shusen, et al.
Published: (2025)
by: Liu, Shusen, et al.
Published: (2025)
MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
by: Wang, Zhiqiang, et al.
Published: (2025)
by: Wang, Zhiqiang, et al.
Published: (2025)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
by: Long, Xiang, et al.
Published: (2026)
by: Long, Xiang, et al.
Published: (2026)
REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks
by: Geng, Longling, et al.
Published: (2025)
by: Geng, Longling, et al.
Published: (2025)
MCP-RADAR: A Multi-Dimensional Benchmark for Evaluating Tool Use Capabilities in Large Language Models
by: Gao, Xuanqi, et al.
Published: (2025)
by: Gao, Xuanqi, et al.
Published: (2025)
Benchmarking LLM Tool-Use in the Wild
by: Yu, Peijie, et al.
Published: (2026)
by: Yu, Peijie, et al.
Published: (2026)
TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
by: He, Pengfei, et al.
Published: (2025)
by: He, Pengfei, et al.
Published: (2025)
When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents
by: Liu, Shi, et al.
Published: (2026)
by: Liu, Shi, et al.
Published: (2026)
MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem
by: Liu, Fan, et al.
Published: (2025)
by: Liu, Fan, et al.
Published: (2025)
OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories
by: Liu, Yibing, et al.
Published: (2026)
by: Liu, Yibing, et al.
Published: (2026)
MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation
by: Yang, Chenghao, et al.
Published: (2025)
by: Yang, Chenghao, et al.
Published: (2025)
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
by: Xu, Haoyuan, et al.
Published: (2026)
by: Xu, Haoyuan, et al.
Published: (2026)
HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks
by: Cui, Fan, et al.
Published: (2026)
by: Cui, Fan, et al.
Published: (2026)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
by: Deng, Shihan, et al.
Published: (2024)
by: Deng, Shihan, et al.
Published: (2024)
AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
by: Shang, Yu, et al.
Published: (2025)
by: Shang, Yu, et al.
Published: (2025)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
by: Li, Yuanyang, et al.
Published: (2026)
by: Li, Yuanyang, et al.
Published: (2026)
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
by: Hu, Lingxiang, et al.
Published: (2026)
by: Hu, Lingxiang, et al.
Published: (2026)
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
by: Xia, Hongfei, et al.
Published: (2025)
by: Xia, Hongfei, et al.
Published: (2025)
TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks
by: Weng, Muyan, et al.
Published: (2026)
by: Weng, Muyan, et al.
Published: (2026)
GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis
by: Yu, Bo, et al.
Published: (2026)
by: Yu, Bo, et al.
Published: (2026)
MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
by: Fei, Xiang, et al.
Published: (2025)
by: Fei, Xiang, et al.
Published: (2025)
OI-Bench: An Option Injection Benchmark for Evaluating LLM Susceptibility to Directive Interference
by: Liou, Yow-Fu, et al.
Published: (2026)
by: Liou, Yow-Fu, et al.
Published: (2026)
$C^3$-Bench: The Things Real Disturbing LLM based Agent in Multi-Tasking
by: Yu, Peijie, et al.
Published: (2025)
by: Yu, Peijie, et al.
Published: (2025)
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
by: Liu, Jiayu, et al.
Published: (2025)
by: Liu, Jiayu, et al.
Published: (2025)
MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments
by: Kong, Quyu, et al.
Published: (2025)
by: Kong, Quyu, et al.
Published: (2025)
MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation
by: Wang, Wenhao, et al.
Published: (2026)
by: Wang, Wenhao, et al.
Published: (2026)
Similar Items
-
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
by: Wang, Zhenting, et al.
Published: (2025) -
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
by: Guo, Zikang, et al.
Published: (2025) -
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
by: Zhu, Jie, et al.
Published: (2026) -
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
by: Jia, Hongrui, et al.
Published: (2025) -
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
by: He, Wei, et al.
Published: (2025)