POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Varela, Iñaki Dellibarda, Sendra-Arranz, R., Romero-Sorozabal, Pablo, Valverde-García, J. M., Laudanski, Annemarie F., Gutiérrez, Álvaro, Rocon, Eduardo, Cebrian, Manuel
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917555628670976
author Varela, Iñaki Dellibarda
Sendra-Arranz, R.
Romero-Sorozabal, Pablo
Valverde-García, J. M.
Laudanski, Annemarie F.
Gutiérrez, Álvaro
Rocon, Eduardo
Cebrian, Manuel
author_facet Varela, Iñaki Dellibarda
Sendra-Arranz, R.
Romero-Sorozabal, Pablo
Valverde-García, J. M.
Laudanski, Annemarie F.
Gutiérrez, Álvaro
Rocon, Eduardo
Cebrian, Manuel
contents Orchestrating Large Language Models into Multi-Agent Systems (LLM-MAS) has unlocked remarkable reasoning capabilities, yet emergent failures and hallucinations that resist characterisation block their deployment in safety-critical domains -- a gap made legally untenable by emerging AI regulation. Existing evaluation paradigms share a common flaw: centralised judgment creates single points of failure and demands domain-specific expertise. Here we present POIROT, a protocol that repurposes a system's own agents as its diagnostic layer, leveraging the epistemic diversity already present in the architecture. Across evaluated settings, POIROT outperforms single-LLM evaluator baselines, with gains that scale with problem complexity (OR = 1.60, $p = 0.008$), agent count, and fault dimensionality, persisting under compound fault conditions. These results demonstrate that safety oversight need not be externalised: the agents executing a role carry sufficient collective intelligence to audit it. We release POIROT as an open-source library alongside BLAME, a benchmark for fault attribution in safety-critical multi-agent systems.
format Preprint
id arxiv_https___arxiv_org_abs_2606_02282
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems
Varela, Iñaki Dellibarda
Sendra-Arranz, R.
Romero-Sorozabal, Pablo
Valverde-García, J. M.
Laudanski, Annemarie F.
Gutiérrez, Álvaro
Rocon, Eduardo
Cebrian, Manuel
Artificial Intelligence
Orchestrating Large Language Models into Multi-Agent Systems (LLM-MAS) has unlocked remarkable reasoning capabilities, yet emergent failures and hallucinations that resist characterisation block their deployment in safety-critical domains -- a gap made legally untenable by emerging AI regulation. Existing evaluation paradigms share a common flaw: centralised judgment creates single points of failure and demands domain-specific expertise. Here we present POIROT, a protocol that repurposes a system's own agents as its diagnostic layer, leveraging the epistemic diversity already present in the architecture. Across evaluated settings, POIROT outperforms single-LLM evaluator baselines, with gains that scale with problem complexity (OR = 1.60, $p = 0.008$), agent count, and fault dimensionality, persisting under compound fault conditions. These results demonstrate that safety oversight need not be externalised: the agents executing a role carry sufficient collective intelligence to audit it. We release POIROT as an open-source library alongside BLAME, a benchmark for fault attribution in safety-critical multi-agent systems.
title POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems
topic Artificial Intelligence
url https://arxiv.org/abs/2606.02282