Fine-Tuning LLMs on Small Medical Datasets: Text Classification and Normalization Effectiveness on Cardiology reports and Discharge records

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Losch, Noah, Plagwitz, Lucas, Büscher, Antonius, Varghese, Julian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909555154223104
author Losch, Noah
Plagwitz, Lucas
Büscher, Antonius
Varghese, Julian
author_facet Losch, Noah
Plagwitz, Lucas
Büscher, Antonius
Varghese, Julian
contents We investigate the effectiveness of fine-tuning large language models (LLMs) on small medical datasets for text classification and named entity recognition tasks. Using a German cardiology report dataset and the i2b2 Smoking Challenge dataset, we demonstrate that fine-tuning small LLMs locally on limited training data can improve performance achieving comparable results to larger models. Our experiments show that fine-tuning improves performance on both tasks, with notable gains observed with as few as 200-300 training examples. Overall, the study highlights the potential of task-specific fine-tuning of LLMs for automating clinical workflows and efficiently extracting structured data from unstructured medical text.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21349
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-Tuning LLMs on Small Medical Datasets: Text Classification and Normalization Effectiveness on Cardiology reports and Discharge records
Losch, Noah
Plagwitz, Lucas
Büscher, Antonius
Varghese, Julian
Computation and Language
Machine Learning
68T50
I.2.6; I.2.7; J.3
We investigate the effectiveness of fine-tuning large language models (LLMs) on small medical datasets for text classification and named entity recognition tasks. Using a German cardiology report dataset and the i2b2 Smoking Challenge dataset, we demonstrate that fine-tuning small LLMs locally on limited training data can improve performance achieving comparable results to larger models. Our experiments show that fine-tuning improves performance on both tasks, with notable gains observed with as few as 200-300 training examples. Overall, the study highlights the potential of task-specific fine-tuning of LLMs for automating clinical workflows and efficiently extracting structured data from unstructured medical text.
title Fine-Tuning LLMs on Small Medical Datasets: Text Classification and Normalization Effectiveness on Cardiology reports and Discharge records
topic Computation and Language
Machine Learning
68T50
I.2.6; I.2.7; J.3
url https://arxiv.org/abs/2503.21349