FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Jie, Tian, Yimin, Li, Boyang, Wu, Kehao, Liang, Zhongzhi, Li, Junhui, Zhang, Xianyin, Guo, Lifan, Chen, Feng, Liu, Yong, Zhang, Chi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
von: Lu, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Lu, Jiaxuan, et al.
Veröffentlicht: (2026)
FinGuard: Detecting Financial Regulatory Non-Compliance in LLM Interactions
von: Dou, Huaixia, et al.
Veröffentlicht: (2026)
von: Dou, Huaixia, et al.
Veröffentlicht: (2026)
MCPToolBench++: A Large Scale AI Agent Model Context Protocol MCP Tool Use Benchmark
von: Fan, Shiqing, et al.
Veröffentlicht: (2025)
von: Fan, Shiqing, et al.
Veröffentlicht: (2025)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
von: Zhang, Dongsen, et al.
Veröffentlicht: (2025)
von: Zhang, Dongsen, et al.
Veröffentlicht: (2025)
Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset
von: Chen, Qian, et al.
Veröffentlicht: (2026)
von: Chen, Qian, et al.
Veröffentlicht: (2026)
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
von: Wang, Zhenting, et al.
Veröffentlicht: (2025)
von: Wang, Zhenting, et al.
Veröffentlicht: (2025)
DianJin-OCR-R1: Enhancing OCR Capabilities via a Reasoning-and-Tool Interleaved Vision-Language Model
von: Chen, Qian, et al.
Veröffentlicht: (2025)
von: Chen, Qian, et al.
Veröffentlicht: (2025)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models
von: Zhu, Jie, et al.
Veröffentlicht: (2025)
von: Zhu, Jie, et al.
Veröffentlicht: (2025)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs
von: Lu, Guilong, et al.
Veröffentlicht: (2025)
von: Lu, Guilong, et al.
Veröffentlicht: (2025)
MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
von: Luo, Ziyang, et al.
Veröffentlicht: (2025)
von: Luo, Ziyang, et al.
Veröffentlicht: (2025)
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
von: Guo, Zikang, et al.
Veröffentlicht: (2025)
von: Guo, Zikang, et al.
Veröffentlicht: (2025)
Benchmarking Large Language Models on CFLUE -- A Chinese Financial Language Understanding Evaluation Dataset
von: Zhu, Jie, et al.
Veröffentlicht: (2024)
von: Zhu, Jie, et al.
Veröffentlicht: (2024)
M$^3$FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset
von: Zhu, Jie, et al.
Veröffentlicht: (2025)
von: Zhu, Jie, et al.
Veröffentlicht: (2025)
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
von: Bandi, Chaithanya, et al.
Veröffentlicht: (2026)
von: Bandi, Chaithanya, et al.
Veröffentlicht: (2026)
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
von: Wang, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Wang, Zhiqiang, et al.
Veröffentlicht: (2025)
MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP Tools
von: Wang, Wenhao, et al.
Veröffentlicht: (2025)
von: Wang, Wenhao, et al.
Veröffentlicht: (2025)
FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain
von: Zeng, Lingfeng, et al.
Veröffentlicht: (2025)
von: Zeng, Lingfeng, et al.
Veröffentlicht: (2025)
MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
von: Zong, Xuanjun, et al.
Veröffentlicht: (2025)
von: Zong, Xuanjun, et al.
Veröffentlicht: (2025)
Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions
von: Hasan, Mohammed Mehedi, et al.
Veröffentlicht: (2026)
von: Hasan, Mohammed Mehedi, et al.
Veröffentlicht: (2026)
ScaleMCP: Dynamic and Auto-Synchronizing Model Context Protocol Tools for LLM Agents
von: Lumer, Elias, et al.
Veröffentlicht: (2025)
von: Lumer, Elias, et al.
Veröffentlicht: (2025)
FinTeam: A Multi-Agent Collaborative Intelligence System for Comprehensive Financial Scenarios
von: Wu, Yingqian, et al.
Veröffentlicht: (2025)
von: Wu, Yingqian, et al.
Veröffentlicht: (2025)
DianJin-R1: Evaluating and Enhancing Financial Reasoning in Large Language Models
von: Zhu, Jie, et al.
Veröffentlicht: (2025)
von: Zhu, Jie, et al.
Veröffentlicht: (2025)
FinTradeBench: A Financial Reasoning Benchmark for LLMs
von: Agrawal, Yogesh, et al.
Veröffentlicht: (2026)
von: Agrawal, Yogesh, et al.
Veröffentlicht: (2026)
FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and Principles
von: Malarkkan, Arun Vignesh, et al.
Veröffentlicht: (2026)
von: Malarkkan, Arun Vignesh, et al.
Veröffentlicht: (2026)
FinSight: Towards Real-World Financial Deep Research
von: Jin, Jiajie, et al.
Veröffentlicht: (2025)
von: Jin, Jiajie, et al.
Veröffentlicht: (2025)
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis
von: Zhang, Qiaohong, et al.
Veröffentlicht: (2026)
von: Zhang, Qiaohong, et al.
Veröffentlicht: (2026)
NetMCP: Network-Aware Model Context Protocol Platform for LLM Capability Extension
von: Li, Enhan, et al.
Veröffentlicht: (2025)
von: Li, Enhan, et al.
Veröffentlicht: (2025)
Benchmarking LLM Tool-Use in the Wild
von: Yu, Peijie, et al.
Veröffentlicht: (2026)
von: Yu, Peijie, et al.
Veröffentlicht: (2026)
Beyond the Protocol: Unveiling Attack Vectors in the Model Context Protocol (MCP) Ecosystem
von: Song, Hao, et al.
Veröffentlicht: (2025)
von: Song, Hao, et al.
Veröffentlicht: (2025)
"MCP Does Not Stand for Misuse Cryptography Protocol": Uncovering Cryptographic Misuse in Model Context Protocol at Scale
von: Yan, Biwei, et al.
Veröffentlicht: (2025)
von: Yan, Biwei, et al.
Veröffentlicht: (2025)
Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers
von: Hasan, Mohammed Mehedi, et al.
Veröffentlicht: (2025)
von: Hasan, Mohammed Mehedi, et al.
Veröffentlicht: (2025)
FinToolSyn: A forward synthesis Framework for Financial Tool-Use Dialogue Data with Dynamic Tool Retrieval
von: Huang, Caishuang, et al.
Veröffentlicht: (2026)
von: Huang, Caishuang, et al.
Veröffentlicht: (2026)
We Urgently Need Privilege Management in MCP: A Measurement of API Usage in MCP Ecosystems
von: Li, Zhihao, et al.
Veröffentlicht: (2025)
von: Li, Zhihao, et al.
Veröffentlicht: (2025)
FinBoardBench: Benchmarking Dynamic Wealth Management and Strategic Financial Reasoning of LLMs via Board Game Simulations
von: Hu, Xuesi, et al.
Veröffentlicht: (2026)
von: Hu, Xuesi, et al.
Veröffentlicht: (2026)
Real Faults in Model Context Protocol (MCP) Software: a Comprehensive Taxonomy
von: Taraghi, Mina, et al.
Veröffentlicht: (2026)
von: Taraghi, Mina, et al.
Veröffentlicht: (2026)
Don't believe everything you read: Understanding and Measuring MCP Behavior under Misleading Tool Descriptions
von: Li, Zhihao, et al.
Veröffentlicht: (2026)
von: Li, Zhihao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
von: Lu, Jiaxuan, et al.
Veröffentlicht: (2026) -
FinGuard: Detecting Financial Regulatory Non-Compliance in LLM Interactions
von: Dou, Huaixia, et al.
Veröffentlicht: (2026) -
MCPToolBench++: A Large Scale AI Agent Model Context Protocol MCP Tool Use Benchmark
von: Fan, Shiqing, et al.
Veröffentlicht: (2025) -
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
von: Zhang, Dongsen, et al.
Veröffentlicht: (2025) -
Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset
von: Chen, Qian, et al.
Veröffentlicht: (2026)