ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Jia-Kai, Huang, I-Wei, Wu, Chun-Tin, Tsai, Yi-Tien |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benchmarking AI for low-resource contexts: Thinking beyond leaderboards
by: Pant, Aakash, et al.
Published: (2026)
by: Pant, Aakash, et al.
Published: (2026)
A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19
by: Khandelwal, Vedant, et al.
Published: (2024)
by: Khandelwal, Vedant, et al.
Published: (2024)
Study Design and Demystification of Physics Informed Neural Networks for Power Flow Simulation
by: Leyli-abadi, Milad, et al.
Published: (2025)
by: Leyli-abadi, Milad, et al.
Published: (2025)
Automated Circuit Interpretation via Probe Prompting
by: Birardi, Giuseppe
Published: (2025)
by: Birardi, Giuseppe
Published: (2025)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
by: Wu, Dekun, et al.
Published: (2023)
by: Wu, Dekun, et al.
Published: (2023)
Augmenting deep neural networks with symbolic knowledge: Towards trustworthy and interpretable AI for education
by: Hooshyar, Danial, et al.
Published: (2023)
by: Hooshyar, Danial, et al.
Published: (2023)
Quantum Abduction: A New Paradigm for Reasoning under Uncertainty
by: Pareschi, Remo
Published: (2025)
by: Pareschi, Remo
Published: (2025)
Gyan: An Explainable Neuro-Symbolic Language Model
by: Srinivasan, Venkat, et al.
Published: (2026)
by: Srinivasan, Venkat, et al.
Published: (2026)
Prescriptive Artificial Intelligence: A Formal Paradigm for Auditing Human Decisions Under Uncertainty
by: Passos, Pedro, et al.
Published: (2025)
by: Passos, Pedro, et al.
Published: (2025)
GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation
by: Ghandi, Taraneh, et al.
Published: (2026)
by: Ghandi, Taraneh, et al.
Published: (2026)
Uncovering Bugs in Formal Explainers: A Case Study with PyXAI
by: Huang, Xuanxiang, et al.
Published: (2025)
by: Huang, Xuanxiang, et al.
Published: (2025)
A Comprehensive Survey of Deep Transfer Learning for Anomaly Detection in Industrial Time Series: Methods, Applications, and Directions
by: Yan, Peng, et al.
Published: (2023)
by: Yan, Peng, et al.
Published: (2023)
OG-RAG: Ontology-Grounded Retrieval-Augmented Generation For Large Language Models
by: Sharma, Kartik, et al.
Published: (2024)
by: Sharma, Kartik, et al.
Published: (2024)
A challenge in A(G)I, cybernetics revived in the Ouroboros Model as one algorithm for all thinking
by: Thomsen, Knud
Published: (2024)
by: Thomsen, Knud
Published: (2024)
RELRaE: LLM-Based Relationship Extraction, Labelling, Refinement, and Evaluation
by: Hannah, George, et al.
Published: (2025)
by: Hannah, George, et al.
Published: (2025)
Evaluating Large Language Models for Causal Modeling
by: Razouk, Houssam, et al.
Published: (2024)
by: Razouk, Houssam, et al.
Published: (2024)
How Metacognitive Architectures Remember Their Own Thoughts: A Systematic Review
by: Nolte, Robin, et al.
Published: (2025)
by: Nolte, Robin, et al.
Published: (2025)
Graph Language Models
by: Plenz, Moritz, et al.
Published: (2024)
by: Plenz, Moritz, et al.
Published: (2024)
AI-generated stories favour stability over change: homogeneity and cultural stereotyping in narratives generated by gpt-4o-mini
by: Rettberg, Jill Walker, et al.
Published: (2025)
by: Rettberg, Jill Walker, et al.
Published: (2025)
Compositional Concept-Based Neuron-Level Interpretability for Deep Reinforcement Learning
by: Jiang, Zeyu, et al.
Published: (2025)
by: Jiang, Zeyu, et al.
Published: (2025)
Development of Hybrid Artificial Intelligence Training on Real and Synthetic Data: Benchmark on Two Mixed Training Strategies
by: Wachter, Paul, et al.
Published: (2025)
by: Wachter, Paul, et al.
Published: (2025)
PRISM-Consult: A Panel-of-Experts Architecture for Clinician-Aligned Diagnosis
by: Levine, Lionel, et al.
Published: (2025)
by: Levine, Lionel, et al.
Published: (2025)
Exploring the Technology Landscape through Topic Modeling, Expert Involvement, and Reinforcement Learning
by: Nazari, Ali, et al.
Published: (2025)
by: Nazari, Ali, et al.
Published: (2025)
5G Traffic Prediction with Time Series Analysis
by: Nayak, Nikhil, et al.
Published: (2021)
by: Nayak, Nikhil, et al.
Published: (2021)
Uniform $\mathcal{C}^k$ Approximation of $G$-Invariant and Antisymmetric Functions, Embedding Dimensions, and Polynomial Representations
by: Ganguly, Soumya, et al.
Published: (2024)
by: Ganguly, Soumya, et al.
Published: (2024)
Explainable Classifier for Malignant Lymphoma Subtyping via Cell Graph and Image Fusion
by: Nishiyama, Daiki, et al.
Published: (2025)
by: Nishiyama, Daiki, et al.
Published: (2025)
Learning Actionable World Models for Industrial Process Control
by: Yan, Peng, et al.
Published: (2025)
by: Yan, Peng, et al.
Published: (2025)
Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems
by: Albiero, Daniel, et al.
Published: (2026)
by: Albiero, Daniel, et al.
Published: (2026)
Making High-Level AI Design Decisions Explicit Using a Binary Stream System-Designation Approach
by: Mossbridge, Julia
Published: (2024)
by: Mossbridge, Julia
Published: (2024)
Machine Learning-Based Detection of MCP Attacks
by: Mattsson, Tobias, et al.
Published: (2026)
by: Mattsson, Tobias, et al.
Published: (2026)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
by: Du, Bangde, et al.
Published: (2025)
by: Du, Bangde, et al.
Published: (2025)
A Human-Machine Collaboration Framework for the Development of Schemas
by: Isaak, Nicos
Published: (2024)
by: Isaak, Nicos
Published: (2024)
SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
by: Nfissi, Alaa, et al.
Published: (2025)
by: Nfissi, Alaa, et al.
Published: (2025)
Boundary Regression for Leitmotif Detection in Music Audio
by: Lee, Sihun, et al.
Published: (2025)
by: Lee, Sihun, et al.
Published: (2025)
A Hierarchical conv-LSTM and LLM Integrated Model for Holistic Stock Forecasting
by: Chakraborty, Arya, et al.
Published: (2024)
by: Chakraborty, Arya, et al.
Published: (2024)
Combination of Weak Learners eXplanations to Improve Random Forest eXplicability Robustness
by: Pala, Riccardo, et al.
Published: (2024)
by: Pala, Riccardo, et al.
Published: (2024)
Attention-based sequential recommendation system using multimodal data
by: Oh, Hyungtaik, et al.
Published: (2024)
by: Oh, Hyungtaik, et al.
Published: (2024)
Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning
by: Li, Lin, et al.
Published: (2026)
by: Li, Lin, et al.
Published: (2026)
A survey of air combat behavior modeling using machine learning
by: Gorton, Patrick Ribu, et al.
Published: (2024)
by: Gorton, Patrick Ribu, et al.
Published: (2024)
ARCTraj: A Dataset and Benchmark of Human Reasoning Trajectories for Abstract Problem Solving
by: Kim, Sejin, et al.
Published: (2025)
by: Kim, Sejin, et al.
Published: (2025)
Similar Items
-
Benchmarking AI for low-resource contexts: Thinking beyond leaderboards
by: Pant, Aakash, et al.
Published: (2026) -
A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19
by: Khandelwal, Vedant, et al.
Published: (2024) -
Study Design and Demystification of Physics Informed Neural Networks for Power Flow Simulation
by: Leyli-abadi, Milad, et al.
Published: (2025) -
Automated Circuit Interpretation via Probe Prompting
by: Birardi, Giuseppe
Published: (2025) -
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
by: Wu, Dekun, et al.
Published: (2023)