Saved in:
Bibliographic Details
Main Authors: Li, Xiang, Yu, Huizi, Wang, Wenkong, Wu, Yiran, Zhou, Jiayan, Hua, Wenyue, Lin, Xinxin, Tan, Wenjia, Zhu, Lexuan, Chen, Bingyi, Chen, Guang, Chen, Ming-Li, Zhou, Yang, Li, Zhao, Assimes, Themistocles L., Zhang, Yongfeng, Wu, Qingyun, Ma, Xin, Li, Lingyao, Fan, Lizhou
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.21228
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908609697284096
author Li, Xiang
Yu, Huizi
Wang, Wenkong
Wu, Yiran
Zhou, Jiayan
Hua, Wenyue
Lin, Xinxin
Tan, Wenjia
Zhu, Lexuan
Chen, Bingyi
Chen, Guang
Chen, Ming-Li
Zhou, Yang
Li, Zhao
Assimes, Themistocles L.
Zhang, Yongfeng
Wu, Qingyun
Ma, Xin
Li, Lingyao
Fan, Lizhou
author_facet Li, Xiang
Yu, Huizi
Wang, Wenkong
Wu, Yiran
Zhou, Jiayan
Hua, Wenyue
Lin, Xinxin
Tan, Wenjia
Zhu, Lexuan
Chen, Bingyi
Chen, Guang
Chen, Ming-Li
Zhou, Yang
Li, Zhao
Assimes, Themistocles L.
Zhang, Yongfeng
Wu, Qingyun
Ma, Xin
Li, Lingyao
Fan, Lizhou
contents Objective: Emergency medical dispatch (EMD) is a high-stakes process challenged by caller distress, ambiguity, and cognitive load. Large Language Models (LLMs) and Multi-Agent Systems (MAS) offer opportunities to augment dispatchers. This study aimed to develop and evaluate a taxonomy-grounded, LLM-powered multi-agent system for simulating realistic EMD scenarios. Methods: We constructed a clinical taxonomy (32 chief complaints, 6 caller identities from MIMIC-III) and a six-phase call protocol. Using this framework, we developed an AutoGen-based MAS with Caller and Dispatcher Agents. The system grounds interactions in a fact commons to ensure clinical plausibility and mitigate misinformation. We used a hybrid evaluation framework: four physicians assessed 100 simulated cases for "Guidance Efficacy" and "Dispatch Effectiveness," supplemented by automated linguistic analysis (sentiment, readability, politeness). Results: Human evaluation, with substantial inter-rater agreement (Gwe's AC1 > 0.70), confirmed the system's high performance. It demonstrated excellent Dispatch Effectiveness (e.g., 94 % contacting the correct potential other agents) and Guidance Efficacy (advice provided in 91 % of cases), both rated highly by physicians. Algorithmic metrics corroborated these findings, indicating a predominantly neutral affective profile (73.7 % neutral sentiment; 90.4 % neutral emotion), high readability (Flesch 80.9), and a consistently polite style (60.0 % polite; 0 % impolite). Conclusion: Our taxonomy-grounded MAS simulates diverse, clinically plausible dispatch scenarios with high fidelity. Findings support its use for dispatcher training, protocol evaluation, and as a foundation for real-time decision support. This work outlines a pathway for safely integrating advanced AI agents into emergency response workflows.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21228
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DispatchMAS: Fusing taxonomy and artificial intelligence agents for emergency medical services
Li, Xiang
Yu, Huizi
Wang, Wenkong
Wu, Yiran
Zhou, Jiayan
Hua, Wenyue
Lin, Xinxin
Tan, Wenjia
Zhu, Lexuan
Chen, Bingyi
Chen, Guang
Chen, Ming-Li
Zhou, Yang
Li, Zhao
Assimes, Themistocles L.
Zhang, Yongfeng
Wu, Qingyun
Ma, Xin
Li, Lingyao
Fan, Lizhou
Computation and Language
Human-Computer Interaction
68T07, 92C50
I.2.7; J.3
Objective: Emergency medical dispatch (EMD) is a high-stakes process challenged by caller distress, ambiguity, and cognitive load. Large Language Models (LLMs) and Multi-Agent Systems (MAS) offer opportunities to augment dispatchers. This study aimed to develop and evaluate a taxonomy-grounded, LLM-powered multi-agent system for simulating realistic EMD scenarios. Methods: We constructed a clinical taxonomy (32 chief complaints, 6 caller identities from MIMIC-III) and a six-phase call protocol. Using this framework, we developed an AutoGen-based MAS with Caller and Dispatcher Agents. The system grounds interactions in a fact commons to ensure clinical plausibility and mitigate misinformation. We used a hybrid evaluation framework: four physicians assessed 100 simulated cases for "Guidance Efficacy" and "Dispatch Effectiveness," supplemented by automated linguistic analysis (sentiment, readability, politeness). Results: Human evaluation, with substantial inter-rater agreement (Gwe's AC1 > 0.70), confirmed the system's high performance. It demonstrated excellent Dispatch Effectiveness (e.g., 94 % contacting the correct potential other agents) and Guidance Efficacy (advice provided in 91 % of cases), both rated highly by physicians. Algorithmic metrics corroborated these findings, indicating a predominantly neutral affective profile (73.7 % neutral sentiment; 90.4 % neutral emotion), high readability (Flesch 80.9), and a consistently polite style (60.0 % polite; 0 % impolite). Conclusion: Our taxonomy-grounded MAS simulates diverse, clinically plausible dispatch scenarios with high fidelity. Findings support its use for dispatcher training, protocol evaluation, and as a foundation for real-time decision support. This work outlines a pathway for safely integrating advanced AI agents into emergency response workflows.
title DispatchMAS: Fusing taxonomy and artificial intelligence agents for emergency medical services
topic Computation and Language
Human-Computer Interaction
68T07, 92C50
I.2.7; J.3
url https://arxiv.org/abs/2510.21228