Evaluating Open-Weight Large Language Models for Structured Data Extraction from Narrative Medical Reports Across Multiple Use Cases and Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Spaanderman, Douwe J., Prathaban, Karthik, Zelina, Petr, Mouheb, Kaouther, Hejtmánek, Lukáš, Marzetti, Matthew, Schurink, Antonius W., Chan, Damian, Niemantsverdriet, Ruben, Hartmann, Frederik, Qian, Zhen, Thomeer, Maarten G. J., Holub, Petr, Akram, Farhan, Wolters, Frank J., Vernooij, Meike W., Verhoef, Cornelis, Bron, Esther E., Nováček, Vít, Grünhagen, Dirk J., Niessen, Wiro J., Starmans, Martijn P. A., Klein, Stefan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914157700317184
author Spaanderman, Douwe J.
Prathaban, Karthik
Zelina, Petr
Mouheb, Kaouther
Hejtmánek, Lukáš
Marzetti, Matthew
Schurink, Antonius W.
Chan, Damian
Niemantsverdriet, Ruben
Hartmann, Frederik
Qian, Zhen
Thomeer, Maarten G. J.
Holub, Petr
Akram, Farhan
Wolters, Frank J.
Vernooij, Meike W.
Verhoef, Cornelis
Bron, Esther E.
Nováček, Vít
Grünhagen, Dirk J.
Niessen, Wiro J.
Starmans, Martijn P. A.
Klein, Stefan
author_facet Spaanderman, Douwe J.
Prathaban, Karthik
Zelina, Petr
Mouheb, Kaouther
Hejtmánek, Lukáš
Marzetti, Matthew
Schurink, Antonius W.
Chan, Damian
Niemantsverdriet, Ruben
Hartmann, Frederik
Qian, Zhen
Thomeer, Maarten G. J.
Holub, Petr
Akram, Farhan
Wolters, Frank J.
Vernooij, Meike W.
Verhoef, Cornelis
Bron, Esther E.
Nováček, Vít
Grünhagen, Dirk J.
Niessen, Wiro J.
Starmans, Martijn P. A.
Klein, Stefan
contents Large language models (LLMs) are increasingly used to extract structured information from free-text clinical records, but prior work often focuses on single tasks, limited models, and English-language reports. We evaluated 15 open-weight LLMs on pathology and radiology reports across six use cases, colorectal liver metastases, liver tumours, neurodegenerative diseases, soft-tissue tumours, melanomas, and sarcomas, at three institutes in the Netherlands, UK, and Czech Republic. Models included general-purpose and medical-specialised LLMs of various sizes, and six prompting strategies were compared: zero-shot, one-shot, few-shot, chain-of-thought, self-consistency, and prompt graph. Performance was assessed using task-appropriate metrics, with consensus rank aggregation and linear mixed-effects models quantifying variance. Top-ranked models achieved macro-average scores close to inter-rater agreement across tasks. Small-to-medium general-purpose models performed comparably to large models, while tiny and specialised models performed worse. Prompt graph and few-shot prompting improved performance by ~13%. Task-specific factors, including variable complexity and annotation variability, influenced results more than model size or prompting strategy. These findings show that open-weight LLMs can extract structured data from clinical reports across diseases, languages, and institutions, offering a scalable approach for clinical data curation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10658
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Open-Weight Large Language Models for Structured Data Extraction from Narrative Medical Reports Across Multiple Use Cases and Languages
Spaanderman, Douwe J.
Prathaban, Karthik
Zelina, Petr
Mouheb, Kaouther
Hejtmánek, Lukáš
Marzetti, Matthew
Schurink, Antonius W.
Chan, Damian
Niemantsverdriet, Ruben
Hartmann, Frederik
Qian, Zhen
Thomeer, Maarten G. J.
Holub, Petr
Akram, Farhan
Wolters, Frank J.
Vernooij, Meike W.
Verhoef, Cornelis
Bron, Esther E.
Nováček, Vít
Grünhagen, Dirk J.
Niessen, Wiro J.
Starmans, Martijn P. A.
Klein, Stefan
Computation and Language
Artificial Intelligence
Large language models (LLMs) are increasingly used to extract structured information from free-text clinical records, but prior work often focuses on single tasks, limited models, and English-language reports. We evaluated 15 open-weight LLMs on pathology and radiology reports across six use cases, colorectal liver metastases, liver tumours, neurodegenerative diseases, soft-tissue tumours, melanomas, and sarcomas, at three institutes in the Netherlands, UK, and Czech Republic. Models included general-purpose and medical-specialised LLMs of various sizes, and six prompting strategies were compared: zero-shot, one-shot, few-shot, chain-of-thought, self-consistency, and prompt graph. Performance was assessed using task-appropriate metrics, with consensus rank aggregation and linear mixed-effects models quantifying variance. Top-ranked models achieved macro-average scores close to inter-rater agreement across tasks. Small-to-medium general-purpose models performed comparably to large models, while tiny and specialised models performed worse. Prompt graph and few-shot prompting improved performance by ~13%. Task-specific factors, including variable complexity and annotation variability, influenced results more than model size or prompting strategy. These findings show that open-weight LLMs can extract structured data from clinical reports across diseases, languages, and institutions, offering a scalable approach for clinical data curation.
title Evaluating Open-Weight Large Language Models for Structured Data Extraction from Narrative Medical Reports Across Multiple Use Cases and Languages
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.10658