LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yin, Ming, Shen, Dinghan, Xu, Silei, Dong, Sixun, Zhang, Mian, Hu, Yebowen, Liu, Shujian, Han, Jianbing, Ma, Simin, Wang, Song, Indurthi, Sathish Reddy, Wang, Xun, Chen, Yiran, Song, Kaiqiang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911709530161152
author Yin, Ming
Shen, Dinghan
Xu, Silei
Dong, Sixun
Zhang, Mian
Hu, Yebowen
Liu, Shujian
Han, Jianbing
Ma, Simin
Wang, Song
Indurthi, Sathish Reddy
Wang, Xun
Chen, Yiran
Song, Kaiqiang
author_facet Yin, Ming
Shen, Dinghan
Xu, Silei
Dong, Sixun
Zhang, Mian
Hu, Yebowen
Liu, Shujian
Han, Jianbing
Ma, Simin
Wang, Song
Indurthi, Sathish Reddy
Wang, Xun
Chen, Yiran
Song, Kaiqiang
contents Tool calling has emerged as a critical capability for AI agents. In contrast to conventional tool calling frameworks that rely on static, provider-specific tool definitions, the Model Context Protocol (MCP) offers a unified interface to discover and invoke tools dynamically. However, there is a significant gap in benchmarking multi-step tasks using diverse MCP tools in realistic, dynamic scenarios. In this work, we present LiveMCP-101, a benchmark of 101 real-world queries that require coordinated use of multiple MCP tools. To address temporal variability in real-world tool responses, we introduce a parallel evaluation framework where a reference agent executes a validated plan simultaneously to produce real-time reference outputs. Experiments show that even frontier LLMs achieve a success rate below 60\%, highlighting challenges in multi-step tool use. Comprehensive error analysis identifies seven failure modes spanning tool planning, parameterization, and output handling, pointing to concrete directions for improving current models. LiveMCP-101 sets a rigorous standard for evaluating real-world agent capabilities, advancing toward autonomous agent systems that reliably execute complex tasks through MCP tool orchestration.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15760
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
Yin, Ming
Shen, Dinghan
Xu, Silei
Dong, Sixun
Zhang, Mian
Hu, Yebowen
Liu, Shujian
Han, Jianbing
Ma, Simin
Wang, Song
Indurthi, Sathish Reddy
Wang, Xun
Chen, Yiran
Song, Kaiqiang
Computation and Language
Artificial Intelligence
Tool calling has emerged as a critical capability for AI agents. In contrast to conventional tool calling frameworks that rely on static, provider-specific tool definitions, the Model Context Protocol (MCP) offers a unified interface to discover and invoke tools dynamically. However, there is a significant gap in benchmarking multi-step tasks using diverse MCP tools in realistic, dynamic scenarios. In this work, we present LiveMCP-101, a benchmark of 101 real-world queries that require coordinated use of multiple MCP tools. To address temporal variability in real-world tool responses, we introduce a parallel evaluation framework where a reference agent executes a validated plan simultaneously to produce real-time reference outputs. Experiments show that even frontier LLMs achieve a success rate below 60\%, highlighting challenges in multi-step tool use. Comprehensive error analysis identifies seven failure modes spanning tool planning, parameterization, and output handling, pointing to concrete directions for improving current models. LiveMCP-101 sets a rigorous standard for evaluating real-world agent capabilities, advancing toward autonomous agent systems that reliably execute complex tasks through MCP tool orchestration.
title LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.15760