Saved in:
Bibliographic Details
Main Authors: Zhang, Yuanyun, Zhang, Mingxuan, Li, Siyuan, Wang, Zihan, Chen, Haoran, Zhou, Wenbo, Li, Shi
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.07706
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908819780534272
author Zhang, Yuanyun
Zhang, Mingxuan
Li, Siyuan
Wang, Zihan
Chen, Haoran
Zhou, Wenbo
Li, Shi
author_facet Zhang, Yuanyun
Zhang, Mingxuan
Li, Siyuan
Wang, Zihan
Chen, Haoran
Zhou, Wenbo
Li, Shi
contents Deep learning models for medical data are typically trained using task specific objectives that encourage representations to collapse onto a small number of discriminative directions. While effective for individual prediction problems, this paradigm underutilizes the rich structure of clinical data and limits the transferability, stability, and interpretability of learned features. In this work, we propose dense feature learning, a representation centric framework that explicitly shapes the linear structure of medical embeddings. Our approach operates directly on embedding matrices, encouraging spectral balance, subspace consistency, and feature orthogonality through objectives defined entirely in terms of linear algebraic properties. Without relying on labels or generative reconstruction, dense feature learning produces representations with higher effective rank, improved conditioning, and greater stability across time. Empirical evaluations across longitudinal EHR data, clinical text, and multimodal patient representations demonstrate consistent improvements in downstream linear performance, robustness, and subspace alignment compared to supervised and self supervised baselines. These results suggest that learning to span clinical variation may be as important as learning to predict clinical outcomes, and position representation geometry as a first class objective in medical AI.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07706
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Dense Feature Learning via Linear Structure Preservation in Medical Data
Zhang, Yuanyun
Zhang, Mingxuan
Li, Siyuan
Wang, Zihan
Chen, Haoran
Zhou, Wenbo
Li, Shi
Machine Learning
Deep learning models for medical data are typically trained using task specific objectives that encourage representations to collapse onto a small number of discriminative directions. While effective for individual prediction problems, this paradigm underutilizes the rich structure of clinical data and limits the transferability, stability, and interpretability of learned features. In this work, we propose dense feature learning, a representation centric framework that explicitly shapes the linear structure of medical embeddings. Our approach operates directly on embedding matrices, encouraging spectral balance, subspace consistency, and feature orthogonality through objectives defined entirely in terms of linear algebraic properties. Without relying on labels or generative reconstruction, dense feature learning produces representations with higher effective rank, improved conditioning, and greater stability across time. Empirical evaluations across longitudinal EHR data, clinical text, and multimodal patient representations demonstrate consistent improvements in downstream linear performance, robustness, and subspace alignment compared to supervised and self supervised baselines. These results suggest that learning to span clinical variation may be as important as learning to predict clinical outcomes, and position representation geometry as a first class objective in medical AI.
title Dense Feature Learning via Linear Structure Preservation in Medical Data
topic Machine Learning
url https://arxiv.org/abs/2602.07706