MEDFuse: Multimodal EHR Data Fusion with Masked Lab-Test Modeling and Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Phan, Thao Minh Nguyen, Dao, Cong-Tinh, Wu, Chenwei, Wang, Jian-Zhe, Liu, Shun, Ding, Jun-En, Restrepo, David, Liu, Feng, Hung, Fang-Ming, Peng, Wen-Chih
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911957878046720
author Phan, Thao Minh Nguyen
Dao, Cong-Tinh
Wu, Chenwei
Wang, Jian-Zhe
Liu, Shun
Ding, Jun-En
Restrepo, David
Liu, Feng
Hung, Fang-Ming
Peng, Wen-Chih
author_facet Phan, Thao Minh Nguyen
Dao, Cong-Tinh
Wu, Chenwei
Wang, Jian-Zhe
Liu, Shun
Ding, Jun-En
Restrepo, David
Liu, Feng
Hung, Fang-Ming
Peng, Wen-Chih
contents Electronic health records (EHRs) are multimodal by nature, consisting of structured tabular features like lab tests and unstructured clinical notes. In real-life clinical practice, doctors use complementary multimodal EHR data sources to get a clearer picture of patients' health and support clinical decision-making. However, most EHR predictive models do not reflect these procedures, as they either focus on a single modality or overlook the inter-modality interactions/redundancy. In this work, we propose MEDFuse, a Multimodal EHR Data Fusion framework that incorporates masked lab-test modeling and large language models (LLMs) to effectively integrate structured and unstructured medical data. MEDFuse leverages multimodal embeddings extracted from two sources: LLMs fine-tuned on free clinical text and masked tabular transformers trained on structured lab test results. We design a disentangled transformer module, optimized by a mutual information loss to 1) decouple modality-specific and modality-shared information and 2) extract useful joint representation from the noise and redundancy present in clinical notes. Through comprehensive validation on the public MIMIC-III dataset and the in-house FEMH dataset, MEDFuse demonstrates great potential in advancing clinical predictions, achieving over 90% F1 score in the 10-disease multi-label classification task.
format Preprint
id arxiv_https___arxiv_org_abs_2407_12309
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MEDFuse: Multimodal EHR Data Fusion with Masked Lab-Test Modeling and Large Language Models
Phan, Thao Minh Nguyen
Dao, Cong-Tinh
Wu, Chenwei
Wang, Jian-Zhe
Liu, Shun
Ding, Jun-En
Restrepo, David
Liu, Feng
Hung, Fang-Ming
Peng, Wen-Chih
Computation and Language
Electronic health records (EHRs) are multimodal by nature, consisting of structured tabular features like lab tests and unstructured clinical notes. In real-life clinical practice, doctors use complementary multimodal EHR data sources to get a clearer picture of patients' health and support clinical decision-making. However, most EHR predictive models do not reflect these procedures, as they either focus on a single modality or overlook the inter-modality interactions/redundancy. In this work, we propose MEDFuse, a Multimodal EHR Data Fusion framework that incorporates masked lab-test modeling and large language models (LLMs) to effectively integrate structured and unstructured medical data. MEDFuse leverages multimodal embeddings extracted from two sources: LLMs fine-tuned on free clinical text and masked tabular transformers trained on structured lab test results. We design a disentangled transformer module, optimized by a mutual information loss to 1) decouple modality-specific and modality-shared information and 2) extract useful joint representation from the noise and redundancy present in clinical notes. Through comprehensive validation on the public MIMIC-III dataset and the in-house FEMH dataset, MEDFuse demonstrates great potential in advancing clinical predictions, achieving over 90% F1 score in the 10-disease multi-label classification task.
title MEDFuse: Multimodal EHR Data Fusion with Masked Lab-Test Modeling and Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2407.12309