COVID-BLUeS -- A Prospective Study on the Value of AI in Lung Ultrasound Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wiedemann, Nina, Boer, Dianne de Korte-de, Richter, Matthias, van de Weijer, Sjors, Buhre, Charlotte, Eggert, Franz A. M., Aarnoudse, Sophie, Grevendonk, Lotte, Röber, Steffen, Remie, Carlijn M. E., Buhre, Wolfgang, Henry, Ronald, Born, Jannis
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916948712882176
author Wiedemann, Nina
Boer, Dianne de Korte-de
Richter, Matthias
van de Weijer, Sjors
Buhre, Charlotte
Eggert, Franz A. M.
Aarnoudse, Sophie
Grevendonk, Lotte
Röber, Steffen
Remie, Carlijn M. E.
Buhre, Wolfgang
Henry, Ronald
Born, Jannis
author_facet Wiedemann, Nina
Boer, Dianne de Korte-de
Richter, Matthias
van de Weijer, Sjors
Buhre, Charlotte
Eggert, Franz A. M.
Aarnoudse, Sophie
Grevendonk, Lotte
Röber, Steffen
Remie, Carlijn M. E.
Buhre, Wolfgang
Henry, Ronald
Born, Jannis
contents As a lightweight and non-invasive imaging technique, lung ultrasound (LUS) has gained importance for assessing lung pathologies. The use of Artificial intelligence (AI) in medical decision support systems is promising due to the time- and expertise-intensive interpretation, however, due to the poor quality of existing data used for training AI models, their usability for real-world applications remains unclear. In a prospective study, we analyze data from 63 COVID-19 suspects (33 positive) collected at Maastricht University Medical Centre. Ultrasound recordings at six body locations were acquired following the BLUE protocol and manually labeled for severity of lung involvement. Several AI models were applied and trained for detection and severity of pulmonary infection. The severity of the lung infection, as assigned by human annotators based on the LUS videos, is not significantly different between COVID-19 positive and negative patients (p = 0.89). Nevertheless, the predictions of image-based AI models identify a COVID-19 infection with 65% accuracy when applied zero-shot (i.e., trained on other datasets), and up to 79% with targeted training, whereas the accuracy based on human annotations is at most 65%. Multi-modal models combining images and CBC improve significantly over image-only models. Although our analysis generally supports the value of AI in LUS assessment, the evaluated models fall short of the performance expected from previous work. We find this is due to 1) the heterogeneity of LUS datasets, limiting the generalization ability to new data, 2) the frame-based processing of AI models ignoring video-level information, and 3) lack of work on multi-modal models that can extract the most relevant information from video-, image- and variable-based inputs. To aid future research, we publish the dataset at: https://github.com/NinaWie/COVID-BLUES.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10556
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle COVID-BLUeS -- A Prospective Study on the Value of AI in Lung Ultrasound Analysis
Wiedemann, Nina
Boer, Dianne de Korte-de
Richter, Matthias
van de Weijer, Sjors
Buhre, Charlotte
Eggert, Franz A. M.
Aarnoudse, Sophie
Grevendonk, Lotte
Röber, Steffen
Remie, Carlijn M. E.
Buhre, Wolfgang
Henry, Ronald
Born, Jannis
Tissues and Organs
Computational Engineering, Finance, and Science
As a lightweight and non-invasive imaging technique, lung ultrasound (LUS) has gained importance for assessing lung pathologies. The use of Artificial intelligence (AI) in medical decision support systems is promising due to the time- and expertise-intensive interpretation, however, due to the poor quality of existing data used for training AI models, their usability for real-world applications remains unclear. In a prospective study, we analyze data from 63 COVID-19 suspects (33 positive) collected at Maastricht University Medical Centre. Ultrasound recordings at six body locations were acquired following the BLUE protocol and manually labeled for severity of lung involvement. Several AI models were applied and trained for detection and severity of pulmonary infection. The severity of the lung infection, as assigned by human annotators based on the LUS videos, is not significantly different between COVID-19 positive and negative patients (p = 0.89). Nevertheless, the predictions of image-based AI models identify a COVID-19 infection with 65% accuracy when applied zero-shot (i.e., trained on other datasets), and up to 79% with targeted training, whereas the accuracy based on human annotations is at most 65%. Multi-modal models combining images and CBC improve significantly over image-only models. Although our analysis generally supports the value of AI in LUS assessment, the evaluated models fall short of the performance expected from previous work. We find this is due to 1) the heterogeneity of LUS datasets, limiting the generalization ability to new data, 2) the frame-based processing of AI models ignoring video-level information, and 3) lack of work on multi-modal models that can extract the most relevant information from video-, image- and variable-based inputs. To aid future research, we publish the dataset at: https://github.com/NinaWie/COVID-BLUES.
title COVID-BLUeS -- A Prospective Study on the Value of AI in Lung Ultrasound Analysis
topic Tissues and Organs
Computational Engineering, Finance, and Science
url https://arxiv.org/abs/2509.10556