Neu-RadBERT for Enhanced Diagnosis of Brain Injuries and Conditions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Singh, Manpreet, Macrae, Sean, Williams, Pierre-Marc, Hung, Nicole, de Franca, Sabrina Araujo, Letourneau-Guillon, Laurent, Carrier, François-Martin, Liu, Bang, Cavayas, Yiorgos Alexandros
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911196931686400
author Singh, Manpreet
Macrae, Sean
Williams, Pierre-Marc
Hung, Nicole
de Franca, Sabrina Araujo
Letourneau-Guillon, Laurent
Carrier, François-Martin
Liu, Bang
Cavayas, Yiorgos Alexandros
author_facet Singh, Manpreet
Macrae, Sean
Williams, Pierre-Marc
Hung, Nicole
de Franca, Sabrina Araujo
Letourneau-Guillon, Laurent
Carrier, François-Martin
Liu, Bang
Cavayas, Yiorgos Alexandros
contents Objective: We sought to develop a classification algorithm to extract diagnoses from free-text radiology reports of brain imaging performed in patients with acute respiratory failure (ARF) undergoing invasive mechanical ventilation. Methods: We developed and fine-tuned Neu-RadBERT, a BERT-based model, to classify unstructured radiology reports. We extracted all the brain imaging reports (computed tomography and magnetic resonance imaging) from MIMIC-IV database, performed in patients with ARF. Initial manual labelling was performed on a subset of reports for various brain abnormalities, followed by fine-tuning Neu-RadBERT using three strategies: 1) baseline RadBERT, 2) Neu-RadBERT with Masked Language Modeling (MLM) pretraining, and 3) Neu-RadBERT with MLM pretraining and oversampling to address data skewness. We compared the performance of this model to Llama-2-13B, an autoregressive LLM. Results: The Neu-RadBERT model, particularly with oversampling, demonstrated significant improvements in diagnostic accuracy compared to baseline RadBERT for brain abnormalities, achieving up to 98.0% accuracy for acute brain injuries. Llama-2-13B exhibited relatively lower performance, peaking at 67.5% binary classification accuracy. This result highlights potential limitations of current autoregressive LLMs for this specific classification task, though it remains possible that larger models or further fine-tuning could improve performance. Conclusion: Neu-RadBERT, enhanced through target domain pretraining and oversampling techniques, offered a robust tool for accurate and reliable diagnosis of neurological conditions from radiology reports. This study underscores the potential of transformer-based NLP models in automatically extracting diagnoses from free text reports with potential applications to both research and patient care.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06232
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Neu-RadBERT for Enhanced Diagnosis of Brain Injuries and Conditions
Singh, Manpreet
Macrae, Sean
Williams, Pierre-Marc
Hung, Nicole
de Franca, Sabrina Araujo
Letourneau-Guillon, Laurent
Carrier, François-Martin
Liu, Bang
Cavayas, Yiorgos Alexandros
Tissues and Organs
Machine Learning
Objective: We sought to develop a classification algorithm to extract diagnoses from free-text radiology reports of brain imaging performed in patients with acute respiratory failure (ARF) undergoing invasive mechanical ventilation. Methods: We developed and fine-tuned Neu-RadBERT, a BERT-based model, to classify unstructured radiology reports. We extracted all the brain imaging reports (computed tomography and magnetic resonance imaging) from MIMIC-IV database, performed in patients with ARF. Initial manual labelling was performed on a subset of reports for various brain abnormalities, followed by fine-tuning Neu-RadBERT using three strategies: 1) baseline RadBERT, 2) Neu-RadBERT with Masked Language Modeling (MLM) pretraining, and 3) Neu-RadBERT with MLM pretraining and oversampling to address data skewness. We compared the performance of this model to Llama-2-13B, an autoregressive LLM. Results: The Neu-RadBERT model, particularly with oversampling, demonstrated significant improvements in diagnostic accuracy compared to baseline RadBERT for brain abnormalities, achieving up to 98.0% accuracy for acute brain injuries. Llama-2-13B exhibited relatively lower performance, peaking at 67.5% binary classification accuracy. This result highlights potential limitations of current autoregressive LLMs for this specific classification task, though it remains possible that larger models or further fine-tuning could improve performance. Conclusion: Neu-RadBERT, enhanced through target domain pretraining and oversampling techniques, offered a robust tool for accurate and reliable diagnosis of neurological conditions from radiology reports. This study underscores the potential of transformer-based NLP models in automatically extracting diagnoses from free text reports with potential applications to both research and patient care.
title Neu-RadBERT for Enhanced Diagnosis of Brain Injuries and Conditions
topic Tissues and Organs
Machine Learning
url https://arxiv.org/abs/2510.06232