PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools
Fuente:
arXiv
Salvato in:
| Autori principali: | Feng, Tianjun, Chen, Yunfeng, Tsai, Chun-Yi, Sun, Yihan, Das, Ayan, Maghraoui, Kaoutar El, Lin, Shuxin, Patel, Dhaval |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Internalizing Tool Knowledge in Small Language Models via QLoRA Fine-Tuning
di: Shemla, Yuval, et al.
Pubblicazione: (2026)
di: Shemla, Yuval, et al.
Pubblicazione: (2026)
MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments
di: Ganapavarapu, Giridhar, et al.
Pubblicazione: (2026)
di: Ganapavarapu, Giridhar, et al.
Pubblicazione: (2026)
Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines
di: Merchant, Alimurtaza Mustafa, et al.
Pubblicazione: (2026)
di: Merchant, Alimurtaza Mustafa, et al.
Pubblicazione: (2026)
Fine-Tuned Thoughts: Leveraging Chain-of-Thought Reasoning for Industrial Asset Health Monitoring
di: Lin, Shuxin, et al.
Pubblicazione: (2025)
di: Lin, Shuxin, et al.
Pubblicazione: (2025)
Towards Automated Solution Recipe Generation for Industrial Asset Management with LLM
di: Zhou, Nianjun, et al.
Pubblicazione: (2024)
di: Zhou, Nianjun, et al.
Pubblicazione: (2024)
Towards Building General Purpose Embedding Models for Industry 4.0 Agents
di: Constantinides, Christodoulos, et al.
Pubblicazione: (2025)
di: Constantinides, Christodoulos, et al.
Pubblicazione: (2025)
SPIN: Structural LLM Planning via Iterative Navigation for Industrial Tasks
di: Ozaki, Yusuke, et al.
Pubblicazione: (2026)
di: Ozaki, Yusuke, et al.
Pubblicazione: (2026)
Analog In-Memory Computing with Uncertainty Quantification for Efficient Edge-based Medical Imaging Segmentation
di: Hamzaoui, Imane, et al.
Pubblicazione: (2024)
di: Hamzaoui, Imane, et al.
Pubblicazione: (2024)
Chat-of-Thought: Collaborative Multi-Agent System for Generating Domain Specific Information
di: Constantinides, Christodoulos, et al.
Pubblicazione: (2025)
di: Constantinides, Christodoulos, et al.
Pubblicazione: (2025)
On the Convergence Theory of Pipeline Gradient-based Analog In-memory Training
di: Wu, Zhaoxian, et al.
Pubblicazione: (2024)
di: Wu, Zhaoxian, et al.
Pubblicazione: (2024)
SparseST: Exploiting Data Sparsity in Spatiotemporal Modeling and Prediction
di: Wu, Junfeng, et al.
Pubblicazione: (2025)
di: Wu, Junfeng, et al.
Pubblicazione: (2025)
Code-Guided Reasoning for Small Language Models: Evaluating Executable MCQA Scaffolds
di: Biswas, Prateek, et al.
Pubblicazione: (2026)
di: Biswas, Prateek, et al.
Pubblicazione: (2026)
Robust Heterogeneous Analog-Digital Computing for Mixture-of-Experts Models with Theoretical Generalization Guarantees
di: Chowdhury, Mohammed Nowaz Rabbani, et al.
Pubblicazione: (2026)
di: Chowdhury, Mohammed Nowaz Rabbani, et al.
Pubblicazione: (2026)
ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem
di: Dong, Jia-Kai, et al.
Pubblicazione: (2025)
di: Dong, Jia-Kai, et al.
Pubblicazione: (2025)
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
di: Wang, Zhenting, et al.
Pubblicazione: (2025)
di: Wang, Zhenting, et al.
Pubblicazione: (2025)
MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents
di: Li, Haonan, et al.
Pubblicazione: (2026)
di: Li, Haonan, et al.
Pubblicazione: (2026)
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
di: Guo, Zikang, et al.
Pubblicazione: (2025)
di: Guo, Zikang, et al.
Pubblicazione: (2025)
LLM Assisted Anomaly Detection Service for Site Reliability Engineers: Enhancing Cloud Infrastructure Resilience
di: Jha, Nimesh, et al.
Pubblicazione: (2025)
di: Jha, Nimesh, et al.
Pubblicazione: (2025)
Paged Attention Meets FlexAttention: Unlocking Long-Context Efficiency in Deployed Inference
di: Joshi, Thomas, et al.
Pubblicazione: (2025)
di: Joshi, Thomas, et al.
Pubblicazione: (2025)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
di: Li, Yuanyang, et al.
Pubblicazione: (2026)
di: Li, Yuanyang, et al.
Pubblicazione: (2026)
MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP Tools
di: Wang, Wenhao, et al.
Pubblicazione: (2025)
di: Wang, Wenhao, et al.
Pubblicazione: (2025)
MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
di: Fei, Xiang, et al.
Pubblicazione: (2025)
di: Fei, Xiang, et al.
Pubblicazione: (2025)
Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees
di: Chowdhury, Mohammed Nowaz Rabbani, et al.
Pubblicazione: (2026)
di: Chowdhury, Mohammed Nowaz Rabbani, et al.
Pubblicazione: (2026)
AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance
di: Patel, Dhaval, et al.
Pubblicazione: (2025)
di: Patel, Dhaval, et al.
Pubblicazione: (2025)
IndustryAssetEQA: A Neurosymbolic Operational Intelligence System for Embodied Question Answering in Industrial Asset Maintenance
di: Shyalika, Chathurangi, et al.
Pubblicazione: (2026)
di: Shyalika, Chathurangi, et al.
Pubblicazione: (2026)
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
di: Jia, Hongrui, et al.
Pubblicazione: (2025)
di: Jia, Hongrui, et al.
Pubblicazione: (2025)
SPIRAL: Symbolic LLM Planning via Grounded and Reflective Search
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
di: Liu, Wenrui, et al.
Pubblicazione: (2025)
di: Liu, Wenrui, et al.
Pubblicazione: (2025)
Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
di: Fan, Zehao, et al.
Pubblicazione: (2025)
di: Fan, Zehao, et al.
Pubblicazione: (2025)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
di: Fang, Yunhua, et al.
Pubblicazione: (2025)
di: Fang, Yunhua, et al.
Pubblicazione: (2025)
VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity LLM with Native MCP Tool Integration
di: Salas Santillana, Juan
Pubblicazione: (2026)
di: Salas Santillana, Juan
Pubblicazione: (2026)
Challenges in Grounding Language in the Real World
di: Lindes, Peter, et al.
Pubblicazione: (2025)
di: Lindes, Peter, et al.
Pubblicazione: (2025)
From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
di: Yue, Ling, et al.
Pubblicazione: (2026)
di: Yue, Ling, et al.
Pubblicazione: (2026)
DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules
di: De Silva, Devin Yasith, et al.
Pubblicazione: (2026)
di: De Silva, Devin Yasith, et al.
Pubblicazione: (2026)
Analog Foundation Models
di: Büchel, Julian, et al.
Pubblicazione: (2025)
di: Büchel, Julian, et al.
Pubblicazione: (2025)
A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts
di: Chowdhury, Mohammed Nowaz Rabbani, et al.
Pubblicazione: (2024)
di: Chowdhury, Mohammed Nowaz Rabbani, et al.
Pubblicazione: (2024)
ScaleMCP: Dynamic and Auto-Synchronizing Model Context Protocol Tools for LLM Agents
di: Lumer, Elias, et al.
Pubblicazione: (2025)
di: Lumer, Elias, et al.
Pubblicazione: (2025)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
di: Guo, Zhengkang, et al.
Pubblicazione: (2026)
di: Guo, Zhengkang, et al.
Pubblicazione: (2026)
Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions
di: Hasan, Mohammed Mehedi, et al.
Pubblicazione: (2026)
di: Hasan, Mohammed Mehedi, et al.
Pubblicazione: (2026)
HumanMCP: A Human-Like Query Dataset for Evaluating MCP Tool Retrieval Performance
di: Laddha, Shubh, et al.
Pubblicazione: (2025)
di: Laddha, Shubh, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Internalizing Tool Knowledge in Small Language Models via QLoRA Fine-Tuning
di: Shemla, Yuval, et al.
Pubblicazione: (2026) -
MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments
di: Ganapavarapu, Giridhar, et al.
Pubblicazione: (2026) -
Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines
di: Merchant, Alimurtaza Mustafa, et al.
Pubblicazione: (2026) -
Fine-Tuned Thoughts: Leveraging Chain-of-Thought Reasoning for Industrial Asset Health Monitoring
di: Lin, Shuxin, et al.
Pubblicazione: (2025) -
Towards Automated Solution Recipe Generation for Industrial Asset Management with LLM
di: Zhou, Nianjun, et al.
Pubblicazione: (2024)