ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Yuanyang, Yang, Xue, Wang, Longyue, Luo, Weihua, Chen, Hongyang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
por: Fei, Xiang, et al.
Publicado: (2025)
por: Fei, Xiang, et al.
Publicado: (2025)
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
por: Bandi, Chaithanya, et al.
Publicado: (2026)
por: Bandi, Chaithanya, et al.
Publicado: (2026)
Dynamic ReAct: Scalable Tool Selection for Large-Scale MCP Environments
por: Gaurav, Nishant, et al.
Publicado: (2025)
por: Gaurav, Nishant, et al.
Publicado: (2025)
Semantic Tool Discovery for Large Language Models: A Vector-Based Approach to MCP Tool Selection
por: Mudunuri, Sarat, et al.
Publicado: (2026)
por: Mudunuri, Sarat, et al.
Publicado: (2026)
RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation
por: Gan, Tiantian, et al.
Publicado: (2025)
por: Gan, Tiantian, et al.
Publicado: (2025)
Agents in the Sandbox: End-to-End Crash Bug Reproduction for Minecraft
por: Yapağcı, Eray, et al.
Publicado: (2025)
por: Yapağcı, Eray, et al.
Publicado: (2025)
CodeMem: Architecting Reproducible Agents via Dynamic MCP and Procedural Memory
por: Gaurav, Nishant, et al.
Publicado: (2025)
por: Gaurav, Nishant, et al.
Publicado: (2025)
DeltaMCP: Incremental Regeneration via Spec-Aware Transformation for MCP servers
por: Pujara, Aditya, et al.
Publicado: (2026)
por: Pujara, Aditya, et al.
Publicado: (2026)
HarnessAPI: A Skill-First Framework for Unified Streaming APIs and MCP Tools
por: Jose, Edwin
Publicado: (2026)
por: Jose, Edwin
Publicado: (2026)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
por: Lai, Peng, et al.
Publicado: (2026)
por: Lai, Peng, et al.
Publicado: (2026)
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
por: Li, Dawei, et al.
Publicado: (2026)
por: Li, Dawei, et al.
Publicado: (2026)
FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning
por: Ding, Haoran, et al.
Publicado: (2026)
por: Ding, Haoran, et al.
Publicado: (2026)
Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
por: Li, Zeping, et al.
Publicado: (2026)
por: Li, Zeping, et al.
Publicado: (2026)
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
por: He, Jiawei, et al.
Publicado: (2026)
por: He, Jiawei, et al.
Publicado: (2026)
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
por: Yan, Shuo, et al.
Publicado: (2025)
por: Yan, Shuo, et al.
Publicado: (2025)
Towards a Declarative Agentic Layer for Intelligent Agents in MCP-Based Server Ecosystems
por: Rodriguez-Sanchez, Maria Jesus, et al.
Publicado: (2026)
por: Rodriguez-Sanchez, Maria Jesus, et al.
Publicado: (2026)
Correct Code, Vulnerable Dependencies: A Large Scale Measurement Study of LLM-Specified Library Versions
por: Wang, Chengjie, et al.
Publicado: (2026)
por: Wang, Chengjie, et al.
Publicado: (2026)
EvolveTool-Bench: Evaluating the Quality of LLM-Generated Tool Libraries as Software Artifacts
por: Kaliyev, Alibek T., et al.
Publicado: (2026)
por: Kaliyev, Alibek T., et al.
Publicado: (2026)
MCP-SandboxScan: WASM-based Secure Execution and Runtime Analysis for MCP Tools
por: Tan, Zhuoran, et al.
Publicado: (2026)
por: Tan, Zhuoran, et al.
Publicado: (2026)
Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling
por: Trae Research Team, et al.
Publicado: (2025)
por: Trae Research Team, et al.
Publicado: (2025)
Solver-Aided Verification of Policy Compliance in Tool-Augmented LLM Agents
por: Winston, Cailin, et al.
Publicado: (2026)
por: Winston, Cailin, et al.
Publicado: (2026)
Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems
por: Xiong, Qian, et al.
Publicado: (2025)
por: Xiong, Qian, et al.
Publicado: (2025)
Schema First Tool APIs for LLM Agents: A Controlled Study of Tool Misuse, Recovery, and Budgeted Performance
por: Sigdel, Akshey, et al.
Publicado: (2026)
por: Sigdel, Akshey, et al.
Publicado: (2026)
Graph-Based Self-Healing Tool Routing for Cost-Efficient LLM Agents
por: Bholani, Neeraj
Publicado: (2026)
por: Bholani, Neeraj
Publicado: (2026)
ToolFuzz -- Automated Agent Tool Testing
por: Milev, Ivan, et al.
Publicado: (2025)
por: Milev, Ivan, et al.
Publicado: (2025)
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
por: Yuan, Danlong, et al.
Publicado: (2026)
por: Yuan, Danlong, et al.
Publicado: (2026)
An LLM-based Quantitative Framework for Evaluating High-Stealthy Backdoor Risks in OSS Supply Chains
por: Yan, Zihe, et al.
Publicado: (2025)
por: Yan, Zihe, et al.
Publicado: (2025)
Thinking Longer, Not Larger: Enhancing Software Engineering Agents via Scaling Test-Time Compute
por: Ma, Yingwei, et al.
Publicado: (2025)
por: Ma, Yingwei, et al.
Publicado: (2025)
Z-Space: A Multi-Agent Tool Orchestration Framework for Enterprise-Grade LLM Automation
por: He, Qingsong, et al.
Publicado: (2025)
por: He, Qingsong, et al.
Publicado: (2025)
We Urgently Need Privilege Management in MCP: A Measurement of API Usage in MCP Ecosystems
por: Li, Zhihao, et al.
Publicado: (2025)
por: Li, Zhihao, et al.
Publicado: (2025)
MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems
por: Jia, Jin, et al.
Publicado: (2026)
por: Jia, Jin, et al.
Publicado: (2026)
ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
por: Shen, Haiyang, et al.
Publicado: (2024)
por: Shen, Haiyang, et al.
Publicado: (2024)
SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
por: Castellani, Tommaso, et al.
Publicado: (2025)
por: Castellani, Tommaso, et al.
Publicado: (2025)
Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools
por: Son, Ha Min, et al.
Publicado: (2025)
por: Son, Ha Min, et al.
Publicado: (2025)
DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use
por: Chen, Aili, et al.
Publicado: (2026)
por: Chen, Aili, et al.
Publicado: (2026)
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
por: Hu, Ruida, et al.
Publicado: (2026)
por: Hu, Ruida, et al.
Publicado: (2026)
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
por: Lindenbauer, Tobias, et al.
Publicado: (2025)
por: Lindenbauer, Tobias, et al.
Publicado: (2025)
Evaluating Software Development Agents: Patch Patterns, Code Quality, and Issue Complexity in Real-World GitHub Scenarios
por: Chen, Zhi, et al.
Publicado: (2024)
por: Chen, Zhi, et al.
Publicado: (2024)
Understanding the Weakness of Large Language Model Agents within a Complex Android Environment
por: Xing, Mingzhe, et al.
Publicado: (2024)
por: Xing, Mingzhe, et al.
Publicado: (2024)
Breaking the Illusion of Identity in LLM Tooling
por: Miller, Marek
Publicado: (2026)
por: Miller, Marek
Publicado: (2026)
Ejemplares similares
-
MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
por: Fei, Xiang, et al.
Publicado: (2025) -
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
por: Bandi, Chaithanya, et al.
Publicado: (2026) -
Dynamic ReAct: Scalable Tool Selection for Large-Scale MCP Environments
por: Gaurav, Nishant, et al.
Publicado: (2025) -
Semantic Tool Discovery for Large Language Models: A Vector-Based Approach to MCP Tool Selection
por: Mudunuri, Sarat, et al.
Publicado: (2026) -
RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation
por: Gan, Tiantian, et al.
Publicado: (2025)