Small Language Models for Emergency Departments Decision Support: A Benchmark Study

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Zirui, Wu, Jiajun, Teitge, Braden, Holodinsky, Jessalyn, Drew, Steve
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916990711496704
author Wang, Zirui
Wu, Jiajun
Teitge, Braden
Holodinsky, Jessalyn
Drew, Steve
author_facet Wang, Zirui
Wu, Jiajun
Teitge, Braden
Holodinsky, Jessalyn
Drew, Steve
contents Large language models (LLMs) have become increasingly popular in medical domains to assist physicians with a variety of clinical and operational tasks. Given the fast-paced and high-stakes environment of emergency departments (EDs), small language models (SLMs), characterized by a reduction in parameter count compared to LLMs, offer significant potential due to their inherent reasoning capability and efficient performance. This enables SLMs to support physicians by providing timely and accurate information synthesis, thereby improving clinical decision-making and workflow efficiency. In this paper, we present a comprehensive benchmark designed to identify SLMs suited for ED decision support, taking into account both specialized medical expertise and broad general problem-solving capabilities. In our evaluations, we focus on SLMs that have been trained on a mixture of general-domain and medical corpora. A key motivation for emphasizing SLMs is the practical hardware limitations, operational cost constraints, and privacy concerns in the typical real-world deployments. Our benchmark datasets include MedMCQA, MedQA-4Options, and PubMedQA, with the medical abstracts dataset emulating tasks aligned with real ED physicians' daily tasks. Experimental results reveal that general-domain SLMs surprisingly outperform their medically fine-tuned counterparts across these diverse benchmarks for ED. This indicates that for ED, specialized medical fine-tuning of the model may not be required.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04032
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Small Language Models for Emergency Departments Decision Support: A Benchmark Study
Wang, Zirui
Wu, Jiajun
Teitge, Braden
Holodinsky, Jessalyn
Drew, Steve
Computation and Language
Artificial Intelligence
Large language models (LLMs) have become increasingly popular in medical domains to assist physicians with a variety of clinical and operational tasks. Given the fast-paced and high-stakes environment of emergency departments (EDs), small language models (SLMs), characterized by a reduction in parameter count compared to LLMs, offer significant potential due to their inherent reasoning capability and efficient performance. This enables SLMs to support physicians by providing timely and accurate information synthesis, thereby improving clinical decision-making and workflow efficiency. In this paper, we present a comprehensive benchmark designed to identify SLMs suited for ED decision support, taking into account both specialized medical expertise and broad general problem-solving capabilities. In our evaluations, we focus on SLMs that have been trained on a mixture of general-domain and medical corpora. A key motivation for emphasizing SLMs is the practical hardware limitations, operational cost constraints, and privacy concerns in the typical real-world deployments. Our benchmark datasets include MedMCQA, MedQA-4Options, and PubMedQA, with the medical abstracts dataset emulating tasks aligned with real ED physicians' daily tasks. Experimental results reveal that general-domain SLMs surprisingly outperform their medically fine-tuned counterparts across these diverse benchmarks for ED. This indicates that for ED, specialized medical fine-tuning of the model may not be required.
title Small Language Models for Emergency Departments Decision Support: A Benchmark Study
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.04032