A Generalist Model for Diverse Text-Guided Medical Image Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cho, Joseph, Mathur, Mrudang, Zakka, Cyril, Kaur, Dhamanpreet, Leipzig, Matthew, Dalal, Alex, Krishnan, Aravind, Koo, Eubee, Wai, Karen, Zhao, Cindy S., Chaudhari, Akshay, Duda, Matthew, Choi, Ashley, Rahimy, Ehsan, Azzouz, Lyna, Fong, Robyn, Shad, Rohan, Hiesinger, William
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915946375938048
author Cho, Joseph
Mathur, Mrudang
Zakka, Cyril
Kaur, Dhamanpreet
Leipzig, Matthew
Dalal, Alex
Krishnan, Aravind
Koo, Eubee
Wai, Karen
Zhao, Cindy S.
Chaudhari, Akshay
Duda, Matthew
Choi, Ashley
Rahimy, Ehsan
Azzouz, Lyna
Fong, Robyn
Shad, Rohan
Hiesinger, William
author_facet Cho, Joseph
Mathur, Mrudang
Zakka, Cyril
Kaur, Dhamanpreet
Leipzig, Matthew
Dalal, Alex
Krishnan, Aravind
Koo, Eubee
Wai, Karen
Zhao, Cindy S.
Chaudhari, Akshay
Duda, Matthew
Choi, Ashley
Rahimy, Ehsan
Azzouz, Lyna
Fong, Robyn
Shad, Rohan
Hiesinger, William
contents Deep learning algorithms require extensive data to achieve robust performance. However, data availability is often restricted in the medical domain due to patient privacy concerns. Synthetic data presents a possible solution to these challenges. Image generative models have found increasing use for medical applications, but are often task-specific, thus limiting their scalability. Moreover, existing models frequently rely on private datasets for training, which constrain their reproducibility. To address this, we introduce MediSyn: an open-access, generalist, text-guided latent diffusion model capable of generating synthetic images across 6 medical specialties and 10 imaging modalities, while being trained exclusively on publicly available data. Through extensive experimentation, we provide several key contributions. First, we demonstrate that training a generative model on visually diverse medical images does not degrade synthetic image quality. Second, we show that this generalist approach is substantially more computationally efficient than a coordinated suite of task-specific models. Third, we establish that a generalist model can produce realistic, text-aligned synthetic images across visually and medically distinct modalities, as validated by expert physicians. Fourth, we provide empirical evidence that these synthetic images are visually distinct from their corresponding real patient images, alleviating concerns about data memorization in image generative models. Finally, we demonstrate that a generalist model can produce synthetic images that improve classifier performance in data-limited settings across multiple medical specialties. Altogether, our findings highlight the immense potential of generalist image generative models to accelerate algorithmic research and development in medicine.
format Preprint
id arxiv_https___arxiv_org_abs_2405_09806
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Generalist Model for Diverse Text-Guided Medical Image Synthesis
Cho, Joseph
Mathur, Mrudang
Zakka, Cyril
Kaur, Dhamanpreet
Leipzig, Matthew
Dalal, Alex
Krishnan, Aravind
Koo, Eubee
Wai, Karen
Zhao, Cindy S.
Chaudhari, Akshay
Duda, Matthew
Choi, Ashley
Rahimy, Ehsan
Azzouz, Lyna
Fong, Robyn
Shad, Rohan
Hiesinger, William
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Deep learning algorithms require extensive data to achieve robust performance. However, data availability is often restricted in the medical domain due to patient privacy concerns. Synthetic data presents a possible solution to these challenges. Image generative models have found increasing use for medical applications, but are often task-specific, thus limiting their scalability. Moreover, existing models frequently rely on private datasets for training, which constrain their reproducibility. To address this, we introduce MediSyn: an open-access, generalist, text-guided latent diffusion model capable of generating synthetic images across 6 medical specialties and 10 imaging modalities, while being trained exclusively on publicly available data. Through extensive experimentation, we provide several key contributions. First, we demonstrate that training a generative model on visually diverse medical images does not degrade synthetic image quality. Second, we show that this generalist approach is substantially more computationally efficient than a coordinated suite of task-specific models. Third, we establish that a generalist model can produce realistic, text-aligned synthetic images across visually and medically distinct modalities, as validated by expert physicians. Fourth, we provide empirical evidence that these synthetic images are visually distinct from their corresponding real patient images, alleviating concerns about data memorization in image generative models. Finally, we demonstrate that a generalist model can produce synthetic images that improve classifier performance in data-limited settings across multiple medical specialties. Altogether, our findings highlight the immense potential of generalist image generative models to accelerate algorithmic research and development in medicine.
title A Generalist Model for Diverse Text-Guided Medical Image Synthesis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2405.09806