Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Shaokun, Yin, Ming, Zhang, Jieyu, Liu, Jiale, Han, Zhiguang, Zhang, Jingyang, Li, Beibin, Wang, Chi, Wang, Huazheng, Chen, Yiran, Wu, Qingyun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913869399588864
author Zhang, Shaokun
Yin, Ming
Zhang, Jieyu
Liu, Jiale
Han, Zhiguang
Zhang, Jingyang
Li, Beibin
Wang, Chi
Wang, Huazheng
Chen, Yiran
Wu, Qingyun
author_facet Zhang, Shaokun
Yin, Ming
Zhang, Jieyu
Liu, Jiale
Han, Zhiguang
Zhang, Jingyang
Li, Beibin
Wang, Chi
Wang, Huazheng
Chen, Yiran
Wu, Qingyun
contents Failure attribution in LLM multi-agent systems-identifying the agent and step responsible for task failures-provides crucial clues for systems debugging but remains underexplored and labor-intensive. In this paper, we propose and formulate a new research area: automated failure attribution for LLM multi-agent systems. To support this initiative, we introduce the Who&When dataset, comprising extensive failure logs from 127 LLM multi-agent systems with fine-grained annotations linking failures to specific agents and decisive error steps. Using the Who&When, we develop and evaluate three automated failure attribution methods, summarizing their corresponding pros and cons. The best method achieves 53.5% accuracy in identifying failure-responsible agents but only 14.2% in pinpointing failure steps, with some methods performing below random. Even SOTA reasoning models, such as OpenAI o1 and DeepSeek R1, fail to achieve practical usability. These results highlight the task's complexity and the need for further research in this area. Code and dataset are available at https://github.com/mingyin1/Agents_Failure_Attribution
format Preprint
id arxiv_https___arxiv_org_abs_2505_00212
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
Zhang, Shaokun
Yin, Ming
Zhang, Jieyu
Liu, Jiale
Han, Zhiguang
Zhang, Jingyang
Li, Beibin
Wang, Chi
Wang, Huazheng
Chen, Yiran
Wu, Qingyun
Multiagent Systems
Computation and Language
Failure attribution in LLM multi-agent systems-identifying the agent and step responsible for task failures-provides crucial clues for systems debugging but remains underexplored and labor-intensive. In this paper, we propose and formulate a new research area: automated failure attribution for LLM multi-agent systems. To support this initiative, we introduce the Who&When dataset, comprising extensive failure logs from 127 LLM multi-agent systems with fine-grained annotations linking failures to specific agents and decisive error steps. Using the Who&When, we develop and evaluate three automated failure attribution methods, summarizing their corresponding pros and cons. The best method achieves 53.5% accuracy in identifying failure-responsible agents but only 14.2% in pinpointing failure steps, with some methods performing below random. Even SOTA reasoning models, such as OpenAI o1 and DeepSeek R1, fail to achieve practical usability. These results highlight the task's complexity and the need for further research in this area. Code and dataset are available at https://github.com/mingyin1/Agents_Failure_Attribution
title Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
topic Multiagent Systems
Computation and Language
url https://arxiv.org/abs/2505.00212