Multi-modal Masked Siamese Network Improves Chest X-Ray Representation Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Shurrab, Saeed, Guerra-Manzanares, Alejandro, Shamout, Farah E.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929521536532480
author Shurrab, Saeed
Guerra-Manzanares, Alejandro
Shamout, Farah E.
author_facet Shurrab, Saeed
Guerra-Manzanares, Alejandro
Shamout, Farah E.
contents Self-supervised learning methods for medical images primarily rely on the imaging modality during pretraining. While such approaches deliver promising results, they do not leverage associated patient or scan information collected within Electronic Health Records (EHR). Here, we propose to incorporate EHR data during self-supervised pretraining with a Masked Siamese Network (MSN) to enhance the quality of chest X-ray representations. We investigate three types of EHR data, including demographic, scan metadata, and inpatient stay information. We evaluate our approach on three publicly available chest X-ray datasets, MIMIC-CXR, CheXpert, and NIH-14, using two vision transformer (ViT) backbones, specifically ViT-Tiny and ViT-Small. In assessing the quality of the representations via linear evaluation, our proposed method demonstrates significant improvement compared to vanilla MSN and state-of-the-art self-supervised learning baselines. Our work highlights the potential of EHR-enhanced self-supervised pre-training for medical imaging. The code is publicly available at: https://github.com/nyuad-cai/CXR-EHR-MSN
format Preprint
id arxiv_https___arxiv_org_abs_2407_04449
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-modal Masked Siamese Network Improves Chest X-Ray Representation Learning
Shurrab, Saeed
Guerra-Manzanares, Alejandro
Shamout, Farah E.
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Self-supervised learning methods for medical images primarily rely on the imaging modality during pretraining. While such approaches deliver promising results, they do not leverage associated patient or scan information collected within Electronic Health Records (EHR). Here, we propose to incorporate EHR data during self-supervised pretraining with a Masked Siamese Network (MSN) to enhance the quality of chest X-ray representations. We investigate three types of EHR data, including demographic, scan metadata, and inpatient stay information. We evaluate our approach on three publicly available chest X-ray datasets, MIMIC-CXR, CheXpert, and NIH-14, using two vision transformer (ViT) backbones, specifically ViT-Tiny and ViT-Small. In assessing the quality of the representations via linear evaluation, our proposed method demonstrates significant improvement compared to vanilla MSN and state-of-the-art self-supervised learning baselines. Our work highlights the potential of EHR-enhanced self-supervised pre-training for medical imaging. The code is publicly available at: https://github.com/nyuad-cai/CXR-EHR-MSN
title Multi-modal Masked Siamese Network Improves Chest X-Ray Representation Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.04449