Data-Driven Discovery of Feature Groups in Clinical Time Series

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sergeev, Fedor, Burger, Manuel, Leshetkina, Polina, Fortuin, Vincent, Rätsch, Gunnar, Kuznetsova, Rita
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914151076462592
author Sergeev, Fedor
Burger, Manuel
Leshetkina, Polina
Fortuin, Vincent
Rätsch, Gunnar
Kuznetsova, Rita
author_facet Sergeev, Fedor
Burger, Manuel
Leshetkina, Polina
Fortuin, Vincent
Rätsch, Gunnar
Kuznetsova, Rita
contents Clinical time series data are critical for patient monitoring and predictive modeling. These time series are typically multivariate and often comprise hundreds of heterogeneous features from different data sources. The grouping of features based on similarity and relevance to the prediction task has been shown to enhance the performance of deep learning architectures. However, defining these groups a priori using only semantic knowledge is challenging, even for domain experts. To address this, we propose a novel method that learns feature groups by clustering weights of feature-wise embedding layers. This approach seamlessly integrates into standard supervised training and discovers the groups that directly improve downstream performance on clinically relevant tasks. We demonstrate that our method outperforms static clustering approaches on synthetic data and achieves performance comparable to expert-defined groups on real-world medical data. Moreover, the learned feature groups are clinically interpretable, enabling data-driven discovery of task-relevant relationships between variables.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08260
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data-Driven Discovery of Feature Groups in Clinical Time Series
Sergeev, Fedor
Burger, Manuel
Leshetkina, Polina
Fortuin, Vincent
Rätsch, Gunnar
Kuznetsova, Rita
Machine Learning
68T07 (Primary), 62H30, 62M10, 62P10 (Secondary)
I.2.6; I.5.3; J.3; G.3
Clinical time series data are critical for patient monitoring and predictive modeling. These time series are typically multivariate and often comprise hundreds of heterogeneous features from different data sources. The grouping of features based on similarity and relevance to the prediction task has been shown to enhance the performance of deep learning architectures. However, defining these groups a priori using only semantic knowledge is challenging, even for domain experts. To address this, we propose a novel method that learns feature groups by clustering weights of feature-wise embedding layers. This approach seamlessly integrates into standard supervised training and discovers the groups that directly improve downstream performance on clinically relevant tasks. We demonstrate that our method outperforms static clustering approaches on synthetic data and achieves performance comparable to expert-defined groups on real-world medical data. Moreover, the learned feature groups are clinically interpretable, enabling data-driven discovery of task-relevant relationships between variables.
title Data-Driven Discovery of Feature Groups in Clinical Time Series
topic Machine Learning
68T07 (Primary), 62H30, 62M10, 62P10 (Secondary)
I.2.6; I.5.3; J.3; G.3
url https://arxiv.org/abs/2511.08260