Structure is Supervision: Multiview Masked Autoencoders for Radiology
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917379639869440 |
|---|---|
| author | Laguna, Sonia Agostini, Andrea Ryser, Alain Ruiperez-Campillo, Samuel Cannistraci, Irene Vandenhirtz, Moritz Mandt, Stephan Deperrois, Nicolas Nooralahzadeh, Farhad Krauthammer, Michael Sutter, Thomas M. Vogt, Julia E. |
| author_facet | Laguna, Sonia Agostini, Andrea Ryser, Alain Ruiperez-Campillo, Samuel Cannistraci, Irene Vandenhirtz, Moritz Mandt, Stephan Deperrois, Nicolas Nooralahzadeh, Farhad Krauthammer, Michael Sutter, Thomas M. Vogt, Julia E. |
| contents | Building robust medical machine learning systems requires pretraining strategies that exploit the intrinsic structure present in clinical data. We introduce Multiview Masked Autoencoder (MVMAE), a self-supervised framework that leverages the natural multi-view organization of radiology studies to learn view-invariant and disease-relevant representations. MVMAE combines masked image reconstruction with cross-view alignment, transforming clinical redundancy across projections into a powerful self-supervisory signal. We further extend this approach with MVMAE-V2T, which incorporates radiology reports as an auxiliary text-based learning signal to enhance semantic grounding while preserving fully vision-based inference. Evaluated on a downstream disease classification task on three large-scale public datasets, MIMIC-CXR, CheXpert, and PadChest, MVMAE consistently outperforms supervised and vision-language baselines. Furthermore, MVMAE-V2T provides additional gains, particularly in low-label regimes where structured textual supervision is most beneficial. Together, these results establish the importance of structural and textual supervision as complementary paths toward scalable, clinically grounded medical foundation models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_22294 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Structure is Supervision: Multiview Masked Autoencoders for Radiology Laguna, Sonia Agostini, Andrea Ryser, Alain Ruiperez-Campillo, Samuel Cannistraci, Irene Vandenhirtz, Moritz Mandt, Stephan Deperrois, Nicolas Nooralahzadeh, Farhad Krauthammer, Michael Sutter, Thomas M. Vogt, Julia E. Computer Vision and Pattern Recognition Machine Learning Building robust medical machine learning systems requires pretraining strategies that exploit the intrinsic structure present in clinical data. We introduce Multiview Masked Autoencoder (MVMAE), a self-supervised framework that leverages the natural multi-view organization of radiology studies to learn view-invariant and disease-relevant representations. MVMAE combines masked image reconstruction with cross-view alignment, transforming clinical redundancy across projections into a powerful self-supervisory signal. We further extend this approach with MVMAE-V2T, which incorporates radiology reports as an auxiliary text-based learning signal to enhance semantic grounding while preserving fully vision-based inference. Evaluated on a downstream disease classification task on three large-scale public datasets, MIMIC-CXR, CheXpert, and PadChest, MVMAE consistently outperforms supervised and vision-language baselines. Furthermore, MVMAE-V2T provides additional gains, particularly in low-label regimes where structured textual supervision is most beneficial. Together, these results establish the importance of structural and textual supervision as complementary paths toward scalable, clinically grounded medical foundation models. |
| title | Structure is Supervision: Multiview Masked Autoencoders for Radiology |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2511.22294 |