MedVAE: Efficient Automated Interpretation of Medical Images with Large-Scale Generalizable Autoencoders

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Varma, Maya, Kumar, Ashwin, van der Sluijs, Rogier, Ostmeier, Sophie, Blankemeier, Louis, Chambon, Pierre, Bluethgen, Christian, Prince, Jip, Langlotz, Curtis, Chaudhari, Akshay
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912410451836928
author Varma, Maya
Kumar, Ashwin
van der Sluijs, Rogier
Ostmeier, Sophie
Blankemeier, Louis
Chambon, Pierre
Bluethgen, Christian
Prince, Jip
Langlotz, Curtis
Chaudhari, Akshay
author_facet Varma, Maya
Kumar, Ashwin
van der Sluijs, Rogier
Ostmeier, Sophie
Blankemeier, Louis
Chambon, Pierre
Bluethgen, Christian
Prince, Jip
Langlotz, Curtis
Chaudhari, Akshay
contents Medical images are acquired at high resolutions with large fields of view in order to capture fine-grained features necessary for clinical decision-making. Consequently, training deep learning models on medical images can incur large computational costs. In this work, we address the challenge of downsizing medical images in order to improve downstream computational efficiency while preserving clinically-relevant features. We introduce MedVAE, a family of six large-scale 2D and 3D autoencoders capable of encoding medical images as downsized latent representations and decoding latent representations back to high-resolution images. We train MedVAE autoencoders using a novel two-stage training approach with 1,052,730 medical images. Across diverse tasks obtained from 20 medical image datasets, we demonstrate that (1) utilizing MedVAE latent representations in place of high-resolution images when training downstream models can lead to efficiency benefits (up to 70x improvement in throughput) while simultaneously preserving clinically-relevant features and (2) MedVAE can decode latent representations back to high-resolution images with high fidelity. Our work demonstrates that large-scale, generalizable autoencoders can help address critical efficiency challenges in the medical domain. Our code is available at https://github.com/StanfordMIMI/MedVAE.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14753
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedVAE: Efficient Automated Interpretation of Medical Images with Large-Scale Generalizable Autoencoders
Varma, Maya
Kumar, Ashwin
van der Sluijs, Rogier
Ostmeier, Sophie
Blankemeier, Louis
Chambon, Pierre
Bluethgen, Christian
Prince, Jip
Langlotz, Curtis
Chaudhari, Akshay
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Medical images are acquired at high resolutions with large fields of view in order to capture fine-grained features necessary for clinical decision-making. Consequently, training deep learning models on medical images can incur large computational costs. In this work, we address the challenge of downsizing medical images in order to improve downstream computational efficiency while preserving clinically-relevant features. We introduce MedVAE, a family of six large-scale 2D and 3D autoencoders capable of encoding medical images as downsized latent representations and decoding latent representations back to high-resolution images. We train MedVAE autoencoders using a novel two-stage training approach with 1,052,730 medical images. Across diverse tasks obtained from 20 medical image datasets, we demonstrate that (1) utilizing MedVAE latent representations in place of high-resolution images when training downstream models can lead to efficiency benefits (up to 70x improvement in throughput) while simultaneously preserving clinically-relevant features and (2) MedVAE can decode latent representations back to high-resolution images with high fidelity. Our work demonstrates that large-scale, generalizable autoencoders can help address critical efficiency challenges in the medical domain. Our code is available at https://github.com/StanfordMIMI/MedVAE.
title MedVAE: Efficient Automated Interpretation of Medical Images with Large-Scale Generalizable Autoencoders
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.14753