General-Purpose vs. Domain-Adapted Large Language Models for Extraction of Structured Data from Chest Radiology Reports

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dhanaliwala, Ali H., Ghosh, Rikhiya, Karn, Sanjeev Kumar, Ullaskrishnan, Poikavila, Farri, Oladimeji, Comaniciu, Dorin, Kahn, Charles E.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913306538672128
author Dhanaliwala, Ali H.
Ghosh, Rikhiya
Karn, Sanjeev Kumar
Ullaskrishnan, Poikavila
Farri, Oladimeji
Comaniciu, Dorin
Kahn, Charles E.
author_facet Dhanaliwala, Ali H.
Ghosh, Rikhiya
Karn, Sanjeev Kumar
Ullaskrishnan, Poikavila
Farri, Oladimeji
Comaniciu, Dorin
Kahn, Charles E.
contents Radiologists produce unstructured data that can be valuable for clinical care when consumed by information systems. However, variability in style limits usage. Study compares system using domain-adapted language model (RadLing) and general-purpose LLM (GPT-4) in extracting relevant features from chest radiology reports and standardizing them to common data elements (CDEs). Three radiologists annotated a retrospective dataset of 1399 chest XR reports (900 training, 499 test) and mapped to 44 pre-selected relevant CDEs. GPT-4 system was prompted with report, feature set, value set, and dynamic few-shots to extract values and map to CDEs. Output key:value pairs were compared to reference standard at both stages and an identical match was considered TP. F1 score for extraction was 97% for RadLing-based system and 78% for GPT-4 system. F1 score for mapping was 98% for RadLing and 94% for GPT-4; difference was statistically significant (P<.001). RadLing's domain-adapted embeddings were better in feature extraction and its light-weight mapper had better f1 score in CDE assignment. RadLing system also demonstrated higher capabilities in differentiating between absent (99% vs 64%) and unspecified (99% vs 89%). RadLing system's domain-adapted embeddings helped improve performance of GPT-4 system to 92% by giving more relevant few-shot prompts. RadLing system offers operational advantages including local deployment and reduced runtime costs.
format Preprint
id arxiv_https___arxiv_org_abs_2311_17213
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle General-Purpose vs. Domain-Adapted Large Language Models for Extraction of Structured Data from Chest Radiology Reports
Dhanaliwala, Ali H.
Ghosh, Rikhiya
Karn, Sanjeev Kumar
Ullaskrishnan, Poikavila
Farri, Oladimeji
Comaniciu, Dorin
Kahn, Charles E.
Computation and Language
Image and Video Processing
Radiologists produce unstructured data that can be valuable for clinical care when consumed by information systems. However, variability in style limits usage. Study compares system using domain-adapted language model (RadLing) and general-purpose LLM (GPT-4) in extracting relevant features from chest radiology reports and standardizing them to common data elements (CDEs). Three radiologists annotated a retrospective dataset of 1399 chest XR reports (900 training, 499 test) and mapped to 44 pre-selected relevant CDEs. GPT-4 system was prompted with report, feature set, value set, and dynamic few-shots to extract values and map to CDEs. Output key:value pairs were compared to reference standard at both stages and an identical match was considered TP. F1 score for extraction was 97% for RadLing-based system and 78% for GPT-4 system. F1 score for mapping was 98% for RadLing and 94% for GPT-4; difference was statistically significant (P<.001). RadLing's domain-adapted embeddings were better in feature extraction and its light-weight mapper had better f1 score in CDE assignment. RadLing system also demonstrated higher capabilities in differentiating between absent (99% vs 64%) and unspecified (99% vs 89%). RadLing system's domain-adapted embeddings helped improve performance of GPT-4 system to 92% by giving more relevant few-shot prompts. RadLing system offers operational advantages including local deployment and reduced runtime costs.
title General-Purpose vs. Domain-Adapted Large Language Models for Extraction of Structured Data from Chest Radiology Reports
topic Computation and Language
Image and Video Processing
url https://arxiv.org/abs/2311.17213