A Question Answering Based Pipeline for Comprehensive Chinese EHR Information Extraction

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ying, Huaiyuan, Yu, Sheng
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913236237942784
author Ying, Huaiyuan
Yu, Sheng
author_facet Ying, Huaiyuan
Yu, Sheng
contents Electronic health records (EHRs) hold significant value for research and applications. As a new way of information extraction, question answering (QA) can extract more flexible information than conventional methods and is more accessible to clinical researchers, but its progress is impeded by the scarcity of annotated data. In this paper, we propose a novel approach that automatically generates training data for transfer learning of QA models. Our pipeline incorporates a preprocessing module to handle challenges posed by extraction types that are not readily compatible with extractive QA frameworks, including cases with discontinuous answers and many-to-one relationships. The obtained QA model exhibits excellent performance on subtasks of information extraction in EHRs, and it can effectively handle few-shot or zero-shot settings involving yes-no questions. Case studies and ablation studies demonstrate the necessity of each component in our design, and the resulting model is deemed suitable for practical use.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11177
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Question Answering Based Pipeline for Comprehensive Chinese EHR Information Extraction
Ying, Huaiyuan
Yu, Sheng
Computation and Language
Information Retrieval
Electronic health records (EHRs) hold significant value for research and applications. As a new way of information extraction, question answering (QA) can extract more flexible information than conventional methods and is more accessible to clinical researchers, but its progress is impeded by the scarcity of annotated data. In this paper, we propose a novel approach that automatically generates training data for transfer learning of QA models. Our pipeline incorporates a preprocessing module to handle challenges posed by extraction types that are not readily compatible with extractive QA frameworks, including cases with discontinuous answers and many-to-one relationships. The obtained QA model exhibits excellent performance on subtasks of information extraction in EHRs, and it can effectively handle few-shot or zero-shot settings involving yes-no questions. Case studies and ablation studies demonstrate the necessity of each component in our design, and the resulting model is deemed suitable for practical use.
title A Question Answering Based Pipeline for Comprehensive Chinese EHR Information Extraction
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2402.11177