Saved in:
Bibliographic Details
Main Authors: Yoon, Kihyuk, Mao, Lingchao, Chong, Catherine, Schwedt, Todd J., Chiang, Chia-Chun, Li, Jing
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.19661
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918363015413760
author Yoon, Kihyuk
Mao, Lingchao
Chong, Catherine
Schwedt, Todd J.
Chiang, Chia-Chun
Li, Jing
author_facet Yoon, Kihyuk
Mao, Lingchao
Chong, Catherine
Schwedt, Todd J.
Chiang, Chia-Chun
Li, Jing
contents Temporal information in structured electronic health records (EHRs) is often lost in sparse one-hot or count-based representations, while sequence models can be costly and data-hungry. We propose PaReGTA, an LLM-based encoding framework that (i) converts longitudinal EHR events into visit-level templated text with explicit temporal cues, (ii) learns domain-adapted visit embeddings via lightweight contrastive fine-tuning of a sentence-embedding model, and (iii) aggregates visit embeddings into a fixed-dimensional patient representation using hybrid temporal pooling that captures both recency and globally informative visits. Because PaReGTA does not require training from scratch but instead utilizes a pre-trained LLM, it can perform well even in data-limited cohorts. Furthermore, PaReGTA is model-agnostic and can benefit from future EHR-specialized sentence-embedding models. For interpretability, we introduce PaReGTA-RSS (Representation Shift Score), which quantifies clinically defined factor importance by recomputing representations after targeted factor removal and projecting representation shifts through a machine learning model. On 39,088 migraine patients from the All of Us Research Program, PaReGTA outperforms sparse baselines for migraine type classification while deep sequential models were unstable in our cohort.
format Preprint
id arxiv_https___arxiv_org_abs_2602_19661
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PaReGTA: An LLM-based EHR Data Encoding Approach to Capture Temporal Information
Yoon, Kihyuk
Mao, Lingchao
Chong, Catherine
Schwedt, Todd J.
Chiang, Chia-Chun
Li, Jing
Machine Learning
Temporal information in structured electronic health records (EHRs) is often lost in sparse one-hot or count-based representations, while sequence models can be costly and data-hungry. We propose PaReGTA, an LLM-based encoding framework that (i) converts longitudinal EHR events into visit-level templated text with explicit temporal cues, (ii) learns domain-adapted visit embeddings via lightweight contrastive fine-tuning of a sentence-embedding model, and (iii) aggregates visit embeddings into a fixed-dimensional patient representation using hybrid temporal pooling that captures both recency and globally informative visits. Because PaReGTA does not require training from scratch but instead utilizes a pre-trained LLM, it can perform well even in data-limited cohorts. Furthermore, PaReGTA is model-agnostic and can benefit from future EHR-specialized sentence-embedding models. For interpretability, we introduce PaReGTA-RSS (Representation Shift Score), which quantifies clinically defined factor importance by recomputing representations after targeted factor removal and projecting representation shifts through a machine learning model. On 39,088 migraine patients from the All of Us Research Program, PaReGTA outperforms sparse baselines for migraine type classification while deep sequential models were unstable in our cohort.
title PaReGTA: An LLM-based EHR Data Encoding Approach to Capture Temporal Information
topic Machine Learning
url https://arxiv.org/abs/2602.19661