MEDFuse: Multimodal EHR Data Fusion with Masked Lab-Test Modeling and Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911957878046720 |
|---|---|
| author | Phan, Thao Minh Nguyen Dao, Cong-Tinh Wu, Chenwei Wang, Jian-Zhe Liu, Shun Ding, Jun-En Restrepo, David Liu, Feng Hung, Fang-Ming Peng, Wen-Chih |
| author_facet | Phan, Thao Minh Nguyen Dao, Cong-Tinh Wu, Chenwei Wang, Jian-Zhe Liu, Shun Ding, Jun-En Restrepo, David Liu, Feng Hung, Fang-Ming Peng, Wen-Chih |
| contents | Electronic health records (EHRs) are multimodal by nature, consisting of structured tabular features like lab tests and unstructured clinical notes. In real-life clinical practice, doctors use complementary multimodal EHR data sources to get a clearer picture of patients' health and support clinical decision-making. However, most EHR predictive models do not reflect these procedures, as they either focus on a single modality or overlook the inter-modality interactions/redundancy. In this work, we propose MEDFuse, a Multimodal EHR Data Fusion framework that incorporates masked lab-test modeling and large language models (LLMs) to effectively integrate structured and unstructured medical data. MEDFuse leverages multimodal embeddings extracted from two sources: LLMs fine-tuned on free clinical text and masked tabular transformers trained on structured lab test results. We design a disentangled transformer module, optimized by a mutual information loss to 1) decouple modality-specific and modality-shared information and 2) extract useful joint representation from the noise and redundancy present in clinical notes. Through comprehensive validation on the public MIMIC-III dataset and the in-house FEMH dataset, MEDFuse demonstrates great potential in advancing clinical predictions, achieving over 90% F1 score in the 10-disease multi-label classification task. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_12309 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MEDFuse: Multimodal EHR Data Fusion with Masked Lab-Test Modeling and Large Language Models Phan, Thao Minh Nguyen Dao, Cong-Tinh Wu, Chenwei Wang, Jian-Zhe Liu, Shun Ding, Jun-En Restrepo, David Liu, Feng Hung, Fang-Ming Peng, Wen-Chih Computation and Language Electronic health records (EHRs) are multimodal by nature, consisting of structured tabular features like lab tests and unstructured clinical notes. In real-life clinical practice, doctors use complementary multimodal EHR data sources to get a clearer picture of patients' health and support clinical decision-making. However, most EHR predictive models do not reflect these procedures, as they either focus on a single modality or overlook the inter-modality interactions/redundancy. In this work, we propose MEDFuse, a Multimodal EHR Data Fusion framework that incorporates masked lab-test modeling and large language models (LLMs) to effectively integrate structured and unstructured medical data. MEDFuse leverages multimodal embeddings extracted from two sources: LLMs fine-tuned on free clinical text and masked tabular transformers trained on structured lab test results. We design a disentangled transformer module, optimized by a mutual information loss to 1) decouple modality-specific and modality-shared information and 2) extract useful joint representation from the noise and redundancy present in clinical notes. Through comprehensive validation on the public MIMIC-III dataset and the in-house FEMH dataset, MEDFuse demonstrates great potential in advancing clinical predictions, achieving over 90% F1 score in the 10-disease multi-label classification task. |
| title | MEDFuse: Multimodal EHR Data Fusion with Masked Lab-Test Modeling and Large Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2407.12309 |