Expert-Guided Prompting and Retrieval-Augmented Generation for Emergency Medical Service Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ge, Xueren, Murtaza, Sahil, Cortez, Anthony, Alemzadeh, Homa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911275029626880
author Ge, Xueren
Murtaza, Sahil
Cortez, Anthony
Alemzadeh, Homa
author_facet Ge, Xueren
Murtaza, Sahil
Cortez, Anthony
Alemzadeh, Homa
contents Large language models (LLMs) have shown promise in medical question answering, yet they often overlook the domain-specific expertise that professionals depend on, such as the clinical subject areas (e.g., trauma, airway) and the certification level (e.g., EMT, Paramedic). Existing approaches typically apply general-purpose prompting or retrieval strategies without leveraging this structured context, limiting performance in high-stakes settings. We address this gap with EMSQA, an 24.3K-question multiple-choice dataset spanning 10 clinical subject areas and 4 certification levels, accompanied by curated, subject area-aligned knowledge bases (40K documents and 2M tokens). Building on EMSQA, we introduce (i) Expert-CoT, a prompting strategy that conditions chain-of-thought (CoT) reasoning on specific clinical subject area and certification level, and (ii) ExpertRAG, a retrieval-augmented generation pipeline that grounds responses in subject area-aligned documents and real-world patient data. Experiments on 4 LLMs show that Expert-CoT improves up to 2.05% over vanilla CoT prompting. Additionally, combining Expert-CoT with ExpertRAG yields up to a 4.59% accuracy gain over standard RAG baselines. Notably, the 32B expertise-augmented LLMs pass all the computer-adaptive EMS certification simulation exams.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10900
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Expert-Guided Prompting and Retrieval-Augmented Generation for Emergency Medical Service Question Answering
Ge, Xueren
Murtaza, Sahil
Cortez, Anthony
Alemzadeh, Homa
Computation and Language
Artificial Intelligence
Large language models (LLMs) have shown promise in medical question answering, yet they often overlook the domain-specific expertise that professionals depend on, such as the clinical subject areas (e.g., trauma, airway) and the certification level (e.g., EMT, Paramedic). Existing approaches typically apply general-purpose prompting or retrieval strategies without leveraging this structured context, limiting performance in high-stakes settings. We address this gap with EMSQA, an 24.3K-question multiple-choice dataset spanning 10 clinical subject areas and 4 certification levels, accompanied by curated, subject area-aligned knowledge bases (40K documents and 2M tokens). Building on EMSQA, we introduce (i) Expert-CoT, a prompting strategy that conditions chain-of-thought (CoT) reasoning on specific clinical subject area and certification level, and (ii) ExpertRAG, a retrieval-augmented generation pipeline that grounds responses in subject area-aligned documents and real-world patient data. Experiments on 4 LLMs show that Expert-CoT improves up to 2.05% over vanilla CoT prompting. Additionally, combining Expert-CoT with ExpertRAG yields up to a 4.59% accuracy gain over standard RAG baselines. Notably, the 32B expertise-augmented LLMs pass all the computer-adaptive EMS certification simulation exams.
title Expert-Guided Prompting and Retrieval-Augmented Generation for Emergency Medical Service Question Answering
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.10900