The Optimization Paradox in Clinical AI Multi-Agent Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bedi, Suhana, Mlauzi, Iddah, Shin, Daniel, Koyejo, Sanmi, Shah, Nigam H.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909646637236224
author Bedi, Suhana
Mlauzi, Iddah
Shin, Daniel
Koyejo, Sanmi
Shah, Nigam H.
author_facet Bedi, Suhana
Mlauzi, Iddah
Shin, Daniel
Koyejo, Sanmi
Shah, Nigam H.
contents Multi-agent artificial intelligence systems are increasingly deployed in clinical settings, yet the relationship between component-level optimization and system-wide performance remains poorly understood. We evaluated this relationship using 2,400 real patient cases from the MIMIC-CDM dataset across four abdominal pathologies (appendicitis, pancreatitis, cholecystitis, diverticulitis), decomposing clinical diagnosis into information gathering, interpretation, and differential diagnosis. We evaluated single agent systems (one model performing all tasks) against multi-agent systems (specialized models for each task) using comprehensive metrics spanning diagnostic outcomes, process adherence, and cost efficiency. Our results reveal a paradox: while multi-agent systems generally outperformed single agents, the component-optimized or Best of Breed system with superior components and excellent process metrics (85.5% information accuracy) significantly underperformed in diagnostic accuracy (67.7% vs. 77.4% for a top multi-agent system). This finding underscores that successful integration of AI in healthcare requires not just component level optimization but also attention to information flow and compatibility between agents. Our findings highlight the need for end to end system validation rather than relying on component metrics alone.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06574
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Optimization Paradox in Clinical AI Multi-Agent Systems
Bedi, Suhana
Mlauzi, Iddah
Shin, Daniel
Koyejo, Sanmi
Shah, Nigam H.
Artificial Intelligence
Multiagent Systems
Multi-agent artificial intelligence systems are increasingly deployed in clinical settings, yet the relationship between component-level optimization and system-wide performance remains poorly understood. We evaluated this relationship using 2,400 real patient cases from the MIMIC-CDM dataset across four abdominal pathologies (appendicitis, pancreatitis, cholecystitis, diverticulitis), decomposing clinical diagnosis into information gathering, interpretation, and differential diagnosis. We evaluated single agent systems (one model performing all tasks) against multi-agent systems (specialized models for each task) using comprehensive metrics spanning diagnostic outcomes, process adherence, and cost efficiency. Our results reveal a paradox: while multi-agent systems generally outperformed single agents, the component-optimized or Best of Breed system with superior components and excellent process metrics (85.5% information accuracy) significantly underperformed in diagnostic accuracy (67.7% vs. 77.4% for a top multi-agent system). This finding underscores that successful integration of AI in healthcare requires not just component level optimization but also attention to information flow and compatibility between agents. Our findings highlight the need for end to end system validation rather than relying on component metrics alone.
title The Optimization Paradox in Clinical AI Multi-Agent Systems
topic Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2506.06574