When Raw Data Prevails: Are Large Language Model Embeddings Effective in Numerical Data Representation for Medical Machine Learning Applications?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gao, Yanjun, Myers, Skatje, Chen, Shan, Dligach, Dmitriy, Miller, Timothy A, Bitterman, Danielle, Churpek, Matthew, Afshar, Majid
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913509078466560
author Gao, Yanjun
Myers, Skatje
Chen, Shan
Dligach, Dmitriy
Miller, Timothy A
Bitterman, Danielle
Churpek, Matthew
Afshar, Majid
author_facet Gao, Yanjun
Myers, Skatje
Chen, Shan
Dligach, Dmitriy
Miller, Timothy A
Bitterman, Danielle
Churpek, Matthew
Afshar, Majid
contents The introduction of Large Language Models (LLMs) has advanced data representation and analysis, bringing significant progress in their use for medical questions and answering. Despite these advancements, integrating tabular data, especially numerical data pivotal in clinical contexts, into LLM paradigms has not been thoroughly explored. In this study, we examine the effectiveness of vector representations from last hidden states of LLMs for medical diagnostics and prognostics using electronic health record (EHR) data. We compare the performance of these embeddings with that of raw numerical EHR data when used as feature inputs to traditional machine learning (ML) algorithms that excel at tabular data learning, such as eXtreme Gradient Boosting. We focus on instruction-tuned LLMs in a zero-shot setting to represent abnormal physiological data and evaluating their utilities as feature extractors to enhance ML classifiers for predicting diagnoses, length of stay, and mortality. Furthermore, we examine prompt engineering techniques on zero-shot and few-shot LLM embeddings to measure their impact comprehensively. Although findings suggest the raw data features still prevails in medical ML tasks, zero-shot LLM embeddings demonstrate competitive results, suggesting a promising avenue for future research in medical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11854
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle When Raw Data Prevails: Are Large Language Model Embeddings Effective in Numerical Data Representation for Medical Machine Learning Applications?
Gao, Yanjun
Myers, Skatje
Chen, Shan
Dligach, Dmitriy
Miller, Timothy A
Bitterman, Danielle
Churpek, Matthew
Afshar, Majid
Computation and Language
Artificial Intelligence
Machine Learning
The introduction of Large Language Models (LLMs) has advanced data representation and analysis, bringing significant progress in their use for medical questions and answering. Despite these advancements, integrating tabular data, especially numerical data pivotal in clinical contexts, into LLM paradigms has not been thoroughly explored. In this study, we examine the effectiveness of vector representations from last hidden states of LLMs for medical diagnostics and prognostics using electronic health record (EHR) data. We compare the performance of these embeddings with that of raw numerical EHR data when used as feature inputs to traditional machine learning (ML) algorithms that excel at tabular data learning, such as eXtreme Gradient Boosting. We focus on instruction-tuned LLMs in a zero-shot setting to represent abnormal physiological data and evaluating their utilities as feature extractors to enhance ML classifiers for predicting diagnoses, length of stay, and mortality. Furthermore, we examine prompt engineering techniques on zero-shot and few-shot LLM embeddings to measure their impact comprehensively. Although findings suggest the raw data features still prevails in medical ML tasks, zero-shot LLM embeddings demonstrate competitive results, suggesting a promising avenue for future research in medical applications.
title When Raw Data Prevails: Are Large Language Model Embeddings Effective in Numerical Data Representation for Medical Machine Learning Applications?
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2408.11854