A Multimodal Pipeline for Clinical Data Extraction: Applying Vision-Language Models to Scans of Transfusion Reaction Reports

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Schäfer, Henning, Schmidt, Cynthia S., Wutzkowsky, Johannes, Lorek, Kamil, Reinartz, Lea, Rückert, Johannes, Temme, Christian, Böckmann, Britta, Horn, Peter A., Friedrich, Christoph M.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916712330297344
author Schäfer, Henning
Schmidt, Cynthia S.
Wutzkowsky, Johannes
Lorek, Kamil
Reinartz, Lea
Rückert, Johannes
Temme, Christian
Böckmann, Britta
Horn, Peter A.
Friedrich, Christoph M.
author_facet Schäfer, Henning
Schmidt, Cynthia S.
Wutzkowsky, Johannes
Lorek, Kamil
Reinartz, Lea
Rückert, Johannes
Temme, Christian
Böckmann, Britta
Horn, Peter A.
Friedrich, Christoph M.
contents Despite the growing adoption of electronic health records, many processes still rely on paper documents, reflecting the heterogeneous real-world conditions in which healthcare is delivered. The manual transcription process is time-consuming and prone to errors when transferring paper-based data to digital formats. To streamline this workflow, this study presents an open-source pipeline that extracts and categorizes checkbox data from scanned documents. Demonstrated on transfusion reaction reports, the design supports adaptation to other checkbox-rich document types. The proposed method integrates checkbox detection, multilingual optical character recognition (OCR) and multilingual vision-language models (VLMs). The pipeline achieves high precision and recall compared against annually compiled gold-standards from 2017 to 2024. The result is a reduction in administrative workload and accurate regulatory reporting. The open-source availability of this pipeline encourages self-hosted parsing of checkbox forms.
format Preprint
id arxiv_https___arxiv_org_abs_2504_20220
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Multimodal Pipeline for Clinical Data Extraction: Applying Vision-Language Models to Scans of Transfusion Reaction Reports
Schäfer, Henning
Schmidt, Cynthia S.
Wutzkowsky, Johannes
Lorek, Kamil
Reinartz, Lea
Rückert, Johannes
Temme, Christian
Böckmann, Britta
Horn, Peter A.
Friedrich, Christoph M.
Computation and Language
Computer Vision and Pattern Recognition
68T07
I.7.5; I.4.7; I.2.7; H.3.3; J.3
Despite the growing adoption of electronic health records, many processes still rely on paper documents, reflecting the heterogeneous real-world conditions in which healthcare is delivered. The manual transcription process is time-consuming and prone to errors when transferring paper-based data to digital formats. To streamline this workflow, this study presents an open-source pipeline that extracts and categorizes checkbox data from scanned documents. Demonstrated on transfusion reaction reports, the design supports adaptation to other checkbox-rich document types. The proposed method integrates checkbox detection, multilingual optical character recognition (OCR) and multilingual vision-language models (VLMs). The pipeline achieves high precision and recall compared against annually compiled gold-standards from 2017 to 2024. The result is a reduction in administrative workload and accurate regulatory reporting. The open-source availability of this pipeline encourages self-hosted parsing of checkbox forms.
title A Multimodal Pipeline for Clinical Data Extraction: Applying Vision-Language Models to Scans of Transfusion Reaction Reports
topic Computation and Language
Computer Vision and Pattern Recognition
68T07
I.7.5; I.4.7; I.2.7; H.3.3; J.3
url https://arxiv.org/abs/2504.20220