MCP-RADAR: A Multi-Dimensional Benchmark for Evaluating Tool Use Capabilities in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Xuanqi, Xie, Siyi, Zhai, Juan, Ma, Shiqing, Shen, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ASSURE: Metamorphic Testing for AI-powered Browser Extensions
by: Gao, Xuanqi, et al.
Published: (2025)
by: Gao, Xuanqi, et al.
Published: (2025)
Mitigating Stylistic Biases of Machine Translation Systems via Monolingual Corpora Only
by: Gao, Xuanqi, et al.
Published: (2025)
by: Gao, Xuanqi, et al.
Published: (2025)
MCPToolBench++: A Large Scale AI Agent Model Context Protocol MCP Tool Use Benchmark
by: Fan, Shiqing, et al.
Published: (2025)
by: Fan, Shiqing, et al.
Published: (2025)
Efficient DNN-Powered Software with Fair Sparse Models
by: Gao, Xuanqi, et al.
Published: (2024)
by: Gao, Xuanqi, et al.
Published: (2024)
Domain-Specialized Tree of Thought through Plug-and-Play Predictors
by: Gao, Xuanqi, et al.
Published: (2026)
by: Gao, Xuanqi, et al.
Published: (2026)
MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
by: Zong, Xuanjun, et al.
Published: (2025)
by: Zong, Xuanjun, et al.
Published: (2025)
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
by: Bandi, Chaithanya, et al.
Published: (2026)
by: Bandi, Chaithanya, et al.
Published: (2026)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
by: Liu, Wenrui, et al.
Published: (2025)
by: Liu, Wenrui, et al.
Published: (2025)
False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models
by: Jiang, Weipeng, et al.
Published: (2026)
by: Jiang, Weipeng, et al.
Published: (2026)
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
by: Lu, Jiarui, et al.
Published: (2024)
by: Lu, Jiarui, et al.
Published: (2024)
DREAM: Debugging and Repairing AutoML Pipelines
by: Zhang, Xiaoyu, et al.
Published: (2023)
by: Zhang, Xiaoyu, et al.
Published: (2023)
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
by: Jang, Kyochul, et al.
Published: (2025)
by: Jang, Kyochul, et al.
Published: (2025)
The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code Generation
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
RADAR: Accelerating Large Language Model Inference With RL-Based Dynamic Draft Trees
by: Ma, Junjie, et al.
Published: (2025)
by: Ma, Junjie, et al.
Published: (2025)
An Optimizable Suffix Is Worth A Thousand Templates: Efficient Black-box Jailbreaking without Affirmative Phrases via LLM as Optimizer
by: Jiang, Weipeng, et al.
Published: (2024)
by: Jiang, Weipeng, et al.
Published: (2024)
Semantic Tool Discovery for Large Language Models: A Vector-Based Approach to MCP Tool Selection
by: Mudunuri, Sarat, et al.
Published: (2026)
by: Mudunuri, Sarat, et al.
Published: (2026)
MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
by: Luo, Ziyang, et al.
Published: (2025)
by: Luo, Ziyang, et al.
Published: (2025)
RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning
by: Ye, Junjie, et al.
Published: (2024)
by: Ye, Junjie, et al.
Published: (2024)
COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models
by: Meaden, James, et al.
Published: (2025)
by: Meaden, James, et al.
Published: (2025)
ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios
by: Ye, Junjie, et al.
Published: (2024)
by: Ye, Junjie, et al.
Published: (2024)
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
by: Guo, Zikang, et al.
Published: (2025)
by: Guo, Zikang, et al.
Published: (2025)
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
by: Li, Ce, et al.
Published: (2025)
by: Li, Ce, et al.
Published: (2025)
In-Context Reinforcement Learning for Tool Use in Large Language Models
by: Ye, Yaoqi, et al.
Published: (2026)
by: Ye, Yaoqi, et al.
Published: (2026)
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
PCEval: A Benchmark for Evaluating Physical Computing Capabilities of Large Language Models
by: Song, Inpyo, et al.
Published: (2025)
by: Song, Inpyo, et al.
Published: (2025)
CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models
by: Padwal, Vedant
Published: (2026)
by: Padwal, Vedant
Published: (2026)
From Effectiveness to Efficiency: Uncovering Linguistic Bias in Large Language Model-based Code Generation
by: Jiang, Weipeng, et al.
Published: (2024)
by: Jiang, Weipeng, et al.
Published: (2024)
ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research
by: Shen, Hao, et al.
Published: (2026)
by: Shen, Hao, et al.
Published: (2026)
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
by: Li, Yuanyang, et al.
Published: (2026)
by: Li, Yuanyang, et al.
Published: (2026)
Large Language Model Agent for Hyper-Parameter Optimization
by: Liu, Siyi, et al.
Published: (2024)
by: Liu, Siyi, et al.
Published: (2024)
Enterprise Large Language Model Evaluation Benchmark
by: Wang, Liya, et al.
Published: (2025)
by: Wang, Liya, et al.
Published: (2025)
MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use
by: Wu, Zijian, et al.
Published: (2025)
by: Wu, Zijian, et al.
Published: (2025)
On the Use of Large Language Models to Generate Capability Ontologies
by: da Silva, Luis Miguel Vieira, et al.
Published: (2024)
by: da Silva, Luis Miguel Vieira, et al.
Published: (2024)
MCP-ITP: An Automated Framework for Implicit Tool Poisoning in MCP
by: Li, Ruiqi, et al.
Published: (2026)
by: Li, Ruiqi, et al.
Published: (2026)
AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
by: Shi, Jinghang, et al.
Published: (2025)
by: Shi, Jinghang, et al.
Published: (2025)
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
by: Holtermann, Carolin, et al.
Published: (2024)
by: Holtermann, Carolin, et al.
Published: (2024)
ParaView-MCP: An Autonomous Visualization Agent with Direct Tool Use
by: Liu, Shusen, et al.
Published: (2025)
by: Liu, Shusen, et al.
Published: (2025)
ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models
by: Zhang, Yuxiang, et al.
Published: (2024)
by: Zhang, Yuxiang, et al.
Published: (2024)
TRUSTVIS: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models
by: Sun, Ruoyu, et al.
Published: (2025)
by: Sun, Ruoyu, et al.
Published: (2025)
Similar Items
-
ASSURE: Metamorphic Testing for AI-powered Browser Extensions
by: Gao, Xuanqi, et al.
Published: (2025) -
Mitigating Stylistic Biases of Machine Translation Systems via Monolingual Corpora Only
by: Gao, Xuanqi, et al.
Published: (2025) -
MCPToolBench++: A Large Scale AI Agent Model Context Protocol MCP Tool Use Benchmark
by: Fan, Shiqing, et al.
Published: (2025) -
Efficient DNN-Powered Software with Fair Sparse Models
by: Gao, Xuanqi, et al.
Published: (2024) -
Domain-Specialized Tree of Thought through Plug-and-Play Predictors
by: Gao, Xuanqi, et al.
Published: (2026)