Leveraging the Structure of Medical Data for Improved Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908463372697600 |
|---|---|
| author | Agostini, Andrea Laguna, Sonia Ryser, Alain Ruiperez-Campillo, Samuel Vandenhirtz, Moritz Deperrois, Nicolas Nooralahzadeh, Farhad Krauthammer, Michael Sutter, Thomas M. Vogt, Julia E. |
| author_facet | Agostini, Andrea Laguna, Sonia Ryser, Alain Ruiperez-Campillo, Samuel Vandenhirtz, Moritz Deperrois, Nicolas Nooralahzadeh, Farhad Krauthammer, Michael Sutter, Thomas M. Vogt, Julia E. |
| contents | Building generalizable medical AI systems requires pretraining strategies that are data-efficient and domain-aware. Unlike internet-scale corpora, clinical datasets such as MIMIC-CXR offer limited image counts and scarce annotations, but exhibit rich internal structure through multi-view imaging. We propose a self-supervised framework that leverages the inherent structure of medical datasets. Specifically, we treat paired chest X-rays (i.e., frontal and lateral views) as natural positive pairs, learning to reconstruct each view from sparse patches while aligning their latent embeddings. Our method requires no textual supervision and produces informative representations. Evaluated on MIMIC-CXR, we show strong performance compared to supervised objectives and baselines being trained without leveraging structure. This work provides a lightweight, modality-agnostic blueprint for domain-specific pretraining where data is structured but scarce |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_02987 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Leveraging the Structure of Medical Data for Improved Representation Learning Agostini, Andrea Laguna, Sonia Ryser, Alain Ruiperez-Campillo, Samuel Vandenhirtz, Moritz Deperrois, Nicolas Nooralahzadeh, Farhad Krauthammer, Michael Sutter, Thomas M. Vogt, Julia E. Computer Vision and Pattern Recognition Machine Learning Building generalizable medical AI systems requires pretraining strategies that are data-efficient and domain-aware. Unlike internet-scale corpora, clinical datasets such as MIMIC-CXR offer limited image counts and scarce annotations, but exhibit rich internal structure through multi-view imaging. We propose a self-supervised framework that leverages the inherent structure of medical datasets. Specifically, we treat paired chest X-rays (i.e., frontal and lateral views) as natural positive pairs, learning to reconstruct each view from sparse patches while aligning their latent embeddings. Our method requires no textual supervision and produces informative representations. Evaluated on MIMIC-CXR, we show strong performance compared to supervised objectives and baselines being trained without leveraging structure. This work provides a lightweight, modality-agnostic blueprint for domain-specific pretraining where data is structured but scarce |
| title | Leveraging the Structure of Medical Data for Improved Representation Learning |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2507.02987 |