Large Language Multimodal Models for 5-Year Chronic Disease Cohort Prediction Using EHR Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Jun-En, Thao, Phan Nguyen Minh, Peng, Wen-Chih, Wang, Jian-Zhe, Chug, Chun-Cheng, Hsieh, Min-Chen, Tseng, Yun-Chien, Chen, Ling, Luo, Dongsheng, Wang, Chi-Te, Chen, Pei-fu, Liu, Feng, Hung, Fang-Ming
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917763155492864
author Ding, Jun-En
Thao, Phan Nguyen Minh
Peng, Wen-Chih
Wang, Jian-Zhe
Chug, Chun-Cheng
Hsieh, Min-Chen
Tseng, Yun-Chien
Chen, Ling
Luo, Dongsheng
Wang, Chi-Te
Chen, Pei-fu
Liu, Feng
Hung, Fang-Ming
author_facet Ding, Jun-En
Thao, Phan Nguyen Minh
Peng, Wen-Chih
Wang, Jian-Zhe
Chug, Chun-Cheng
Hsieh, Min-Chen
Tseng, Yun-Chien
Chen, Ling
Luo, Dongsheng
Wang, Chi-Te
Chen, Pei-fu
Liu, Feng
Hung, Fang-Ming
contents Chronic diseases such as diabetes are the leading causes of morbidity and mortality worldwide. Numerous research studies have been attempted with various deep learning models in diagnosis. However, most previous studies had certain limitations, including using publicly available datasets (e.g. MIMIC), and imbalanced data. In this study, we collected five-year electronic health records (EHRs) from the Taiwan hospital database, including 1,420,596 clinical notes, 387,392 laboratory test results, and more than 1,505 laboratory test items, focusing on research pre-training large language models. We proposed a novel Large Language Multimodal Models (LLMMs) framework incorporating multimodal data from clinical notes and laboratory test results for the prediction of chronic disease risk. Our method combined a text embedding encoder and multi-head attention layer to learn laboratory test values, utilizing a deep neural network (DNN) module to merge blood features with chronic disease semantics into a latent space. In our experiments, we observe that clinicalBERT and PubMed-BERT, when combined with attention fusion, can achieve an accuracy of 73% in multiclass chronic diseases and diabetes prediction. By transforming laboratory test values into textual descriptions and employing the Flan T-5 model, we achieved a 76% Area Under the ROC Curve (AUROC), demonstrating the effectiveness of leveraging numerical text data for training and inference in language models. This approach significantly improves the accuracy of early-stage diabetes prediction.
format Preprint
id arxiv_https___arxiv_org_abs_2403_04785
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large Language Multimodal Models for 5-Year Chronic Disease Cohort Prediction Using EHR Data
Ding, Jun-En
Thao, Phan Nguyen Minh
Peng, Wen-Chih
Wang, Jian-Zhe
Chug, Chun-Cheng
Hsieh, Min-Chen
Tseng, Yun-Chien
Chen, Ling
Luo, Dongsheng
Wang, Chi-Te
Chen, Pei-fu
Liu, Feng
Hung, Fang-Ming
Computation and Language
Artificial Intelligence
Chronic diseases such as diabetes are the leading causes of morbidity and mortality worldwide. Numerous research studies have been attempted with various deep learning models in diagnosis. However, most previous studies had certain limitations, including using publicly available datasets (e.g. MIMIC), and imbalanced data. In this study, we collected five-year electronic health records (EHRs) from the Taiwan hospital database, including 1,420,596 clinical notes, 387,392 laboratory test results, and more than 1,505 laboratory test items, focusing on research pre-training large language models. We proposed a novel Large Language Multimodal Models (LLMMs) framework incorporating multimodal data from clinical notes and laboratory test results for the prediction of chronic disease risk. Our method combined a text embedding encoder and multi-head attention layer to learn laboratory test values, utilizing a deep neural network (DNN) module to merge blood features with chronic disease semantics into a latent space. In our experiments, we observe that clinicalBERT and PubMed-BERT, when combined with attention fusion, can achieve an accuracy of 73% in multiclass chronic diseases and diabetes prediction. By transforming laboratory test values into textual descriptions and employing the Flan T-5 model, we achieved a 76% Area Under the ROC Curve (AUROC), demonstrating the effectiveness of leveraging numerical text data for training and inference in language models. This approach significantly improves the accuracy of early-stage diabetes prediction.
title Large Language Multimodal Models for 5-Year Chronic Disease Cohort Prediction Using EHR Data
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2403.04785