FirstAidQA: A Synthetic Dataset for First Aid and Emergency Response in Low-Connectivity Settings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Muna, Saiyma Sittul, Salvi, Rezwan Islam, Mushfique, Mushfiqur Rahman, Abrar, Ajwad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912683663556608
author Muna, Saiyma Sittul
Salvi, Rezwan Islam
Mushfique, Mushfiqur Rahman
Abrar, Ajwad
author_facet Muna, Saiyma Sittul
Salvi, Rezwan Islam
Mushfique, Mushfiqur Rahman
Abrar, Ajwad
contents In emergency situations, every second counts. The deployment of Large Language Models (LLMs) in time-sensitive, low or zero-connectivity environments remains limited. Current models are computationally intensive and unsuitable for low-tier devices often used by first responders or civilians. A major barrier to developing lightweight, domain-specific solutions is the lack of high-quality datasets tailored to first aid and emergency response. To address this gap, we introduce FirstAidQA, a synthetic dataset containing 5,500 high-quality question answer pairs that encompass a wide range of first aid and emergency response scenarios. The dataset was generated using a Large Language Model, ChatGPT-4o-mini, with prompt-based in-context learning, using texts from the Vital First Aid Book (2019). We applied preprocessing steps such as text cleaning, contextual chunking, and filtering, followed by human validation to ensure accuracy, safety, and practical relevance of the QA pairs. FirstAidQA is designed to support instruction-tuning and fine-tuning of LLMs and Small Language Models (SLMs), enabling faster, more reliable, and offline-capable systems for emergency settings. We publicly release the dataset to advance research on safety-critical and resource-constrained AI applications in first aid and emergency response. The dataset is available on Hugging Face at https://huggingface.co/datasets/i-am-mushfiq/FirstAidQA.
format Preprint
id arxiv_https___arxiv_org_abs_2511_01289
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FirstAidQA: A Synthetic Dataset for First Aid and Emergency Response in Low-Connectivity Settings
Muna, Saiyma Sittul
Salvi, Rezwan Islam
Mushfique, Mushfiqur Rahman
Abrar, Ajwad
Computation and Language
In emergency situations, every second counts. The deployment of Large Language Models (LLMs) in time-sensitive, low or zero-connectivity environments remains limited. Current models are computationally intensive and unsuitable for low-tier devices often used by first responders or civilians. A major barrier to developing lightweight, domain-specific solutions is the lack of high-quality datasets tailored to first aid and emergency response. To address this gap, we introduce FirstAidQA, a synthetic dataset containing 5,500 high-quality question answer pairs that encompass a wide range of first aid and emergency response scenarios. The dataset was generated using a Large Language Model, ChatGPT-4o-mini, with prompt-based in-context learning, using texts from the Vital First Aid Book (2019). We applied preprocessing steps such as text cleaning, contextual chunking, and filtering, followed by human validation to ensure accuracy, safety, and practical relevance of the QA pairs. FirstAidQA is designed to support instruction-tuning and fine-tuning of LLMs and Small Language Models (SLMs), enabling faster, more reliable, and offline-capable systems for emergency settings. We publicly release the dataset to advance research on safety-critical and resource-constrained AI applications in first aid and emergency response. The dataset is available on Hugging Face at https://huggingface.co/datasets/i-am-mushfiq/FirstAidQA.
title FirstAidQA: A Synthetic Dataset for First Aid and Emergency Response in Low-Connectivity Settings
topic Computation and Language
url https://arxiv.org/abs/2511.01289