Give me Some Hard Questions: Synthetic Data Generation for Clinical QA

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bai, Fan, Harrigian, Keith, Stremmel, Joel, Hassanzadeh, Hamid, Saeedi, Ardavan, Dredze, Mark
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913599468863488
author Bai, Fan
Harrigian, Keith
Stremmel, Joel
Hassanzadeh, Hamid
Saeedi, Ardavan
Dredze, Mark
author_facet Bai, Fan
Harrigian, Keith
Stremmel, Joel
Hassanzadeh, Hamid
Saeedi, Ardavan
Dredze, Mark
contents Clinical Question Answering (QA) systems enable doctors to quickly access patient information from electronic health records (EHRs). However, training these systems requires significant annotated data, which is limited due to the expertise needed and the privacy concerns associated with clinical data. This paper explores generating Clinical QA data using large language models (LLMs) in a zero-shot setting. We find that naive prompting often results in easy questions that do not reflect the complexity of clinical scenarios. To address this, we propose two prompting strategies: 1) instructing the model to generate questions that do not overlap with the input context, and 2) summarizing the input record using a predefined schema to scaffold question generation. Experiments on two Clinical QA datasets demonstrate that our method generates more challenging questions, significantly improving fine-tuning performance over baselines. We compare synthetic and gold data and find a gap between their training efficacy resulting from the quality of synthetically generated answers.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04573
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Give me Some Hard Questions: Synthetic Data Generation for Clinical QA
Bai, Fan
Harrigian, Keith
Stremmel, Joel
Hassanzadeh, Hamid
Saeedi, Ardavan
Dredze, Mark
Computation and Language
Clinical Question Answering (QA) systems enable doctors to quickly access patient information from electronic health records (EHRs). However, training these systems requires significant annotated data, which is limited due to the expertise needed and the privacy concerns associated with clinical data. This paper explores generating Clinical QA data using large language models (LLMs) in a zero-shot setting. We find that naive prompting often results in easy questions that do not reflect the complexity of clinical scenarios. To address this, we propose two prompting strategies: 1) instructing the model to generate questions that do not overlap with the input context, and 2) summarizing the input record using a predefined schema to scaffold question generation. Experiments on two Clinical QA datasets demonstrate that our method generates more challenging questions, significantly improving fine-tuning performance over baselines. We compare synthetic and gold data and find a gap between their training efficacy resulting from the quality of synthetically generated answers.
title Give me Some Hard Questions: Synthetic Data Generation for Clinical QA
topic Computation and Language
url https://arxiv.org/abs/2412.04573