When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Huang, Donghao, Malwe, Gauri, Wang, Zhaoxia |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
par: Huang, Donghao, et autres
Publié: (2025)
par: Huang, Donghao, et autres
Publié: (2025)
Beyond Task Success: Measuring Workflow Fidelity in LLM-Based Agentic Payment Systems
par: Huang, Donghao, et autres
Publié: (2026)
par: Huang, Donghao, et autres
Publié: (2026)
When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail
par: Li, Xiaoxiao
Publié: (2026)
par: Li, Xiaoxiao
Publié: (2026)
Explainable Sentiment Analysis with DeepSeek-R1: Performance, Efficiency, and Few-Shot Learning
par: Huang, Donghao, et autres
Publié: (2025)
par: Huang, Donghao, et autres
Publié: (2025)
Why Do Multi-Agent LLM Systems Fail?
par: Cemri, Mert, et autres
Publié: (2025)
par: Cemri, Mert, et autres
Publié: (2025)
When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems
par: Wang, Zehao, et autres
Publié: (2026)
par: Wang, Zehao, et autres
Publié: (2026)
Task Complexity Matters: An Empirical Study of Reasoning in LLMs for Sentiment Analysis
par: Huang, Donghao, et autres
Publié: (2026)
par: Huang, Donghao, et autres
Publié: (2026)
STARS: Skill-Triggered Audit for Request-Conditioned Invocation Safety in Agent Systems
par: Zhang, Guijia, et autres
Publié: (2026)
par: Zhang, Guijia, et autres
Publié: (2026)
Talk Structurally, Act Hierarchically: A Collaborative Framework for LLM Multi-Agent Systems
par: Wang, Zhao, et autres
Publié: (2025)
par: Wang, Zhao, et autres
Publié: (2025)
Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems
par: Xiong, Qian, et autres
Publié: (2025)
par: Xiong, Qian, et autres
Publié: (2025)
LM Agents May Fail to Act on Their Own Risk Knowledge
par: Tang, Yuzhi, et autres
Publié: (2025)
par: Tang, Yuzhi, et autres
Publié: (2025)
RSAgent: Learning to Reason and Act for Text-Guided Segmentation via Multi-Turn Tool Invocations
par: He, Xingqi, et autres
Publié: (2025)
par: He, Xingqi, et autres
Publié: (2025)
Advancing and Benchmarking Personalized Tool Invocation for LLMs
par: Huang, Xu, et autres
Publié: (2025)
par: Huang, Xu, et autres
Publié: (2025)
Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment
par: Xie, Yuchong, et autres
Publié: (2025)
par: Xie, Yuchong, et autres
Publié: (2025)
Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use
par: Guo, Ruocheng, et autres
Publié: (2026)
par: Guo, Ruocheng, et autres
Publié: (2026)
Towards Unifying Quantitative Security Benchmarking for Multi Agent Systems
par: Sharma, Gauri, et autres
Publié: (2025)
par: Sharma, Gauri, et autres
Publié: (2025)
AgentGit: A Version Control Framework for Reliable and Scalable LLM-Powered Multi-Agent Systems
par: Li, Yang, et autres
Publié: (2025)
par: Li, Yang, et autres
Publié: (2025)
Toward Verifiable Misinformation Detection: A Multi-Tool LLM Agent Framework
par: Cui, Zikun, et autres
Publié: (2025)
par: Cui, Zikun, et autres
Publié: (2025)
Logic Agent: Enhancing Validity with Logic Rule Invocation
par: Liu, Hanmeng, et autres
Publié: (2024)
par: Liu, Hanmeng, et autres
Publié: (2024)
When Does Memory Help Multi-Trajectory Inference for Tool-Use LLM Agents?
par: Li, Xinzhe, et autres
Publié: (2026)
par: Li, Xinzhe, et autres
Publié: (2026)
Zodiac: A Cardiologist-Level LLM Framework for Multi-Agent Diagnostics
par: Zhou, Yuan, et autres
Publié: (2024)
par: Zhou, Yuan, et autres
Publié: (2024)
Gradientsys: A Multi-Agent LLM Scheduler with ReAct Orchestration
par: Song, Xinyuan, et autres
Publié: (2025)
par: Song, Xinyuan, et autres
Publié: (2025)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
par: Hadeliya, Tsimur, et autres
Publié: (2025)
par: Hadeliya, Tsimur, et autres
Publié: (2025)
Automated Creation and Enrichment Framework for Improved Invocation of Enterprise APIs as Tools
par: Agarwal, Prerna, et autres
Publié: (2025)
par: Agarwal, Prerna, et autres
Publié: (2025)
Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution
par: He, Kaiwen, et autres
Publié: (2025)
par: He, Kaiwen, et autres
Publié: (2025)
GeoMind: An Agentic Workflow for Lithology Classification with Reasoned Tool Invocation
par: Zhou, Yitong, et autres
Publié: (2026)
par: Zhou, Yitong, et autres
Publié: (2026)
When LLM Reward Design Fails: Diagnostic-Driven Refinement for Sparse Structured RL
par: Wang, Youting, et autres
Publié: (2026)
par: Wang, Youting, et autres
Publié: (2026)
Why Retrying Fails: Context Contamination in LLM Agent Pipelines
par: Yang, Zhanfu
Publié: (2026)
par: Yang, Zhanfu
Publié: (2026)
HGMF: A Hierarchical Gaussian Mixture Framework for Scalable Tool Invocation within the Model Context Protocol
par: Xing, Wenpeng, et autres
Publié: (2025)
par: Xing, Wenpeng, et autres
Publié: (2025)
Pre-Act: Multi-Step Planning and Reasoning Improves Acting in LLM Agents
par: Rawat, Mrinal, et autres
Publié: (2025)
par: Rawat, Mrinal, et autres
Publié: (2025)
A Flexible Multi-Agent LLM-Human Framework for Fast Human Validated Tool Building
par: Xavier, Daull, et autres
Publié: (2025)
par: Xavier, Daull, et autres
Publié: (2025)
A Dual-Positive Monotone Parameterization for Multi-Segment Bids and a Validity Assessment Framework for Reinforcement Learning Agent-based Simulation of Electricity Markets
par: Xu, Zunnan, et autres
Publié: (2026)
par: Xu, Zunnan, et autres
Publié: (2026)
DroidCall: A Dataset for LLM-powered Android Intent Invocation
par: Xie, Weikai, et autres
Publié: (2024)
par: Xie, Weikai, et autres
Publié: (2024)
When Agents "Misremember" Collectively: Exploring the Mandela Effect in LLM-based Multi-Agent Systems
par: Xu, Naen, et autres
Publié: (2026)
par: Xu, Naen, et autres
Publié: (2026)
MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems
par: Jia, Jin, et autres
Publié: (2026)
par: Jia, Jin, et autres
Publié: (2026)
A Novel Hierarchical Multi-Agent System for Payments Using LLMs
par: Chua, Joon Kiat, et autres
Publié: (2026)
par: Chua, Joon Kiat, et autres
Publié: (2026)
When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents
par: Sheng, Strick, et autres
Publié: (2026)
par: Sheng, Strick, et autres
Publié: (2026)
ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management
par: Pan, Zaifeng, et autres
Publié: (2026)
par: Pan, Zaifeng, et autres
Publié: (2026)
Z-Space: A Multi-Agent Tool Orchestration Framework for Enterprise-Grade LLM Automation
par: He, Qingsong, et autres
Publié: (2025)
par: He, Qingsong, et autres
Publié: (2025)
ReAct Meets ActRe: When Language Agents Enjoy Training Data Autonomy
par: Yang, Zonghan, et autres
Publié: (2024)
par: Yang, Zonghan, et autres
Publié: (2024)
Documents similaires
-
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
par: Huang, Donghao, et autres
Publié: (2025) -
Beyond Task Success: Measuring Workflow Fidelity in LLM-Based Agentic Payment Systems
par: Huang, Donghao, et autres
Publié: (2026) -
When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail
par: Li, Xiaoxiao
Publié: (2026) -
Explainable Sentiment Analysis with DeepSeek-R1: Performance, Efficiency, and Few-Shot Learning
par: Huang, Donghao, et autres
Publié: (2025) -
Why Do Multi-Agent LLM Systems Fail?
par: Cemri, Mert, et autres
Publié: (2025)