Rethinking Tokenization for Clinical Time Series: When Less is More

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Attrach, Rafi Al, Fani, Rajna, Restrepo, David, Jia, Yugang, Schüffler, Peter
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914180411424768
author Attrach, Rafi Al
Fani, Rajna
Restrepo, David
Jia, Yugang
Schüffler, Peter
author_facet Attrach, Rafi Al
Fani, Rajna
Restrepo, David
Jia, Yugang
Schüffler, Peter
contents Tokenization strategies shape how models process electronic health records, yet fair comparisons of their effectiveness remain limited. We present a systematic evaluation of tokenization approaches for clinical time series modeling using transformer-based architectures, revealing task-dependent and sometimes counterintuitive findings about temporal and value feature importance. Through controlled ablations across four clinical prediction tasks on MIMIC-IV, we demonstrate that explicit time encodings provide no consistent statistically significant benefit for the evaluated downstream tasks. Value features show task-dependent importance, affecting mortality prediction but not readmission, suggesting code sequences alone can carry sufficient predictive signal. We further show that frozen pretrained code encoders dramatically outperform their trainable counterparts while requiring dramatically fewer parameters. Larger clinical encoders provide consistent improvements across tasks, benefiting from frozen embeddings that eliminate computational overhead. Our controlled evaluation enables fairer tokenization comparisons and demonstrates that simpler, parameter-efficient approaches can, in many cases, achieve strong performance, though the optimal tokenization strategy remains task-dependent.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05217
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Tokenization for Clinical Time Series: When Less is More
Attrach, Rafi Al
Fani, Rajna
Restrepo, David
Jia, Yugang
Schüffler, Peter
Machine Learning
Tokenization strategies shape how models process electronic health records, yet fair comparisons of their effectiveness remain limited. We present a systematic evaluation of tokenization approaches for clinical time series modeling using transformer-based architectures, revealing task-dependent and sometimes counterintuitive findings about temporal and value feature importance. Through controlled ablations across four clinical prediction tasks on MIMIC-IV, we demonstrate that explicit time encodings provide no consistent statistically significant benefit for the evaluated downstream tasks. Value features show task-dependent importance, affecting mortality prediction but not readmission, suggesting code sequences alone can carry sufficient predictive signal. We further show that frozen pretrained code encoders dramatically outperform their trainable counterparts while requiring dramatically fewer parameters. Larger clinical encoders provide consistent improvements across tasks, benefiting from frozen embeddings that eliminate computational overhead. Our controlled evaluation enables fairer tokenization comparisons and demonstrates that simpler, parameter-efficient approaches can, in many cases, achieve strong performance, though the optimal tokenization strategy remains task-dependent.
title Rethinking Tokenization for Clinical Time Series: When Less is More
topic Machine Learning
url https://arxiv.org/abs/2512.05217