Contextual Phenotyping of Pediatric Sepsis Cohort Using Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nagori, Aditya, Gautam, Ayush, Wiens, Matthew O., Nguyen, Vuong, Mugisha, Nathan Kenya, Kabakyenga, Jerome, Kissoon, Niranjan, Ansermino, John Mark, Kamaleswaran, Rishikesan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918113301233664
author Nagori, Aditya
Gautam, Ayush
Wiens, Matthew O.
Nguyen, Vuong
Mugisha, Nathan Kenya
Kabakyenga, Jerome
Kissoon, Niranjan
Ansermino, John Mark
Kamaleswaran, Rishikesan
author_facet Nagori, Aditya
Gautam, Ayush
Wiens, Matthew O.
Nguyen, Vuong
Mugisha, Nathan Kenya
Kabakyenga, Jerome
Kissoon, Niranjan
Ansermino, John Mark
Kamaleswaran, Rishikesan
contents Clustering patient subgroups is essential for personalized care and efficient resource use. Traditional clustering methods struggle with high-dimensional, heterogeneous healthcare data and lack contextual understanding. This study evaluates Large Language Model (LLM) based clustering against classical methods using a pediatric sepsis dataset from a low-income country (LIC), containing 2,686 records with 28 numerical and 119 categorical variables. Patient records were serialized into text with and without a clustering objective. Embeddings were generated using quantized LLAMA 3.1 8B, DeepSeek-R1-Distill-Llama-8B with low-rank adaptation(LoRA), and Stella-En-400M-V5 models. K-means clustering was applied to these embeddings. Classical comparisons included K-Medoids clustering on UMAP and FAMD-reduced mixed data. Silhouette scores and statistical tests evaluated cluster quality and distinctiveness. Stella-En-400M-V5 achieved the highest Silhouette Score (0.86). LLAMA 3.1 8B with the clustering objective performed better with higher number of clusters, identifying subgroups with distinct nutritional, clinical, and socioeconomic profiles. LLM-based methods outperformed classical techniques by capturing richer context and prioritizing key features. These results highlight potential of LLMs for contextual phenotyping and informed decision-making in resource-limited settings.
format Preprint
id arxiv_https___arxiv_org_abs_2505_09805
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Contextual Phenotyping of Pediatric Sepsis Cohort Using Large Language Models
Nagori, Aditya
Gautam, Ayush
Wiens, Matthew O.
Nguyen, Vuong
Mugisha, Nathan Kenya
Kabakyenga, Jerome
Kissoon, Niranjan
Ansermino, John Mark
Kamaleswaran, Rishikesan
Quantitative Methods
Artificial Intelligence
Computation and Language
Machine Learning
Applications
Clustering patient subgroups is essential for personalized care and efficient resource use. Traditional clustering methods struggle with high-dimensional, heterogeneous healthcare data and lack contextual understanding. This study evaluates Large Language Model (LLM) based clustering against classical methods using a pediatric sepsis dataset from a low-income country (LIC), containing 2,686 records with 28 numerical and 119 categorical variables. Patient records were serialized into text with and without a clustering objective. Embeddings were generated using quantized LLAMA 3.1 8B, DeepSeek-R1-Distill-Llama-8B with low-rank adaptation(LoRA), and Stella-En-400M-V5 models. K-means clustering was applied to these embeddings. Classical comparisons included K-Medoids clustering on UMAP and FAMD-reduced mixed data. Silhouette scores and statistical tests evaluated cluster quality and distinctiveness. Stella-En-400M-V5 achieved the highest Silhouette Score (0.86). LLAMA 3.1 8B with the clustering objective performed better with higher number of clusters, identifying subgroups with distinct nutritional, clinical, and socioeconomic profiles. LLM-based methods outperformed classical techniques by capturing richer context and prioritizing key features. These results highlight potential of LLMs for contextual phenotyping and informed decision-making in resource-limited settings.
title Contextual Phenotyping of Pediatric Sepsis Cohort Using Large Language Models
topic Quantitative Methods
Artificial Intelligence
Computation and Language
Machine Learning
Applications
url https://arxiv.org/abs/2505.09805