Structure is Supervision: Multiview Masked Autoencoders for Radiology

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Laguna, Sonia, Agostini, Andrea, Ryser, Alain, Ruiperez-Campillo, Samuel, Cannistraci, Irene, Vandenhirtz, Moritz, Mandt, Stephan, Deperrois, Nicolas, Nooralahzadeh, Farhad, Krauthammer, Michael, Sutter, Thomas M., Vogt, Julia E.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917379639869440
author Laguna, Sonia
Agostini, Andrea
Ryser, Alain
Ruiperez-Campillo, Samuel
Cannistraci, Irene
Vandenhirtz, Moritz
Mandt, Stephan
Deperrois, Nicolas
Nooralahzadeh, Farhad
Krauthammer, Michael
Sutter, Thomas M.
Vogt, Julia E.
author_facet Laguna, Sonia
Agostini, Andrea
Ryser, Alain
Ruiperez-Campillo, Samuel
Cannistraci, Irene
Vandenhirtz, Moritz
Mandt, Stephan
Deperrois, Nicolas
Nooralahzadeh, Farhad
Krauthammer, Michael
Sutter, Thomas M.
Vogt, Julia E.
contents Building robust medical machine learning systems requires pretraining strategies that exploit the intrinsic structure present in clinical data. We introduce Multiview Masked Autoencoder (MVMAE), a self-supervised framework that leverages the natural multi-view organization of radiology studies to learn view-invariant and disease-relevant representations. MVMAE combines masked image reconstruction with cross-view alignment, transforming clinical redundancy across projections into a powerful self-supervisory signal. We further extend this approach with MVMAE-V2T, which incorporates radiology reports as an auxiliary text-based learning signal to enhance semantic grounding while preserving fully vision-based inference. Evaluated on a downstream disease classification task on three large-scale public datasets, MIMIC-CXR, CheXpert, and PadChest, MVMAE consistently outperforms supervised and vision-language baselines. Furthermore, MVMAE-V2T provides additional gains, particularly in low-label regimes where structured textual supervision is most beneficial. Together, these results establish the importance of structural and textual supervision as complementary paths toward scalable, clinically grounded medical foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22294
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structure is Supervision: Multiview Masked Autoencoders for Radiology
Laguna, Sonia
Agostini, Andrea
Ryser, Alain
Ruiperez-Campillo, Samuel
Cannistraci, Irene
Vandenhirtz, Moritz
Mandt, Stephan
Deperrois, Nicolas
Nooralahzadeh, Farhad
Krauthammer, Michael
Sutter, Thomas M.
Vogt, Julia E.
Computer Vision and Pattern Recognition
Machine Learning
Building robust medical machine learning systems requires pretraining strategies that exploit the intrinsic structure present in clinical data. We introduce Multiview Masked Autoencoder (MVMAE), a self-supervised framework that leverages the natural multi-view organization of radiology studies to learn view-invariant and disease-relevant representations. MVMAE combines masked image reconstruction with cross-view alignment, transforming clinical redundancy across projections into a powerful self-supervisory signal. We further extend this approach with MVMAE-V2T, which incorporates radiology reports as an auxiliary text-based learning signal to enhance semantic grounding while preserving fully vision-based inference. Evaluated on a downstream disease classification task on three large-scale public datasets, MIMIC-CXR, CheXpert, and PadChest, MVMAE consistently outperforms supervised and vision-language baselines. Furthermore, MVMAE-V2T provides additional gains, particularly in low-label regimes where structured textual supervision is most beneficial. Together, these results establish the importance of structural and textual supervision as complementary paths toward scalable, clinically grounded medical foundation models.
title Structure is Supervision: Multiview Masked Autoencoders for Radiology
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2511.22294