MAISI-v2: Accelerated 3D High-Resolution Medical Image Synthesis with Rectified Flow and Region-specific Contrastive Loss

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Can, Guo, Pengfei, Yang, Dong, Tang, Yucheng, He, Yufan, Simon, Benjamin, Belue, Mason, Harmon, Stephanie, Turkbey, Baris, Xu, Daguang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915619646996480
author Zhao, Can
Guo, Pengfei
Yang, Dong
Tang, Yucheng
He, Yufan
Simon, Benjamin
Belue, Mason
Harmon, Stephanie
Turkbey, Baris
Xu, Daguang
author_facet Zhao, Can
Guo, Pengfei
Yang, Dong
Tang, Yucheng
He, Yufan
Simon, Benjamin
Belue, Mason
Harmon, Stephanie
Turkbey, Baris
Xu, Daguang
contents Medical image synthesis is an important topic for both clinical and research applications. Recently, diffusion models have become a leading approach in this area. Despite their strengths, many existing methods struggle with (1) limited generalizability that only work for specific body regions or voxel spacings, (2) slow inference, which is a common issue for diffusion models, and (3) weak alignment with input conditions, which is a critical issue for medical imaging. MAISI, a previously proposed framework, addresses generalizability issues but still suffers from slow inference and limited condition consistency. In this work, we present MAISI-v2, the first accelerated 3D medical image synthesis framework that integrates rectified flow to enable fast and high quality generation. To further enhance condition fidelity, we introduce a novel region-specific contrastive loss to enhance the sensitivity to region of interest. Our experiments show that MAISI-v2 can achieve SOTA image quality with $33 \times$ acceleration for latent diffusion model. We also conducted a downstream segmentation experiment to show that the synthetic images can be used for data augmentation. We release our code, training details, model weights, and a GUI demo to facilitate reproducibility and promote further development within the community.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05772
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MAISI-v2: Accelerated 3D High-Resolution Medical Image Synthesis with Rectified Flow and Region-specific Contrastive Loss
Zhao, Can
Guo, Pengfei
Yang, Dong
Tang, Yucheng
He, Yufan
Simon, Benjamin
Belue, Mason
Harmon, Stephanie
Turkbey, Baris
Xu, Daguang
Computer Vision and Pattern Recognition
Medical image synthesis is an important topic for both clinical and research applications. Recently, diffusion models have become a leading approach in this area. Despite their strengths, many existing methods struggle with (1) limited generalizability that only work for specific body regions or voxel spacings, (2) slow inference, which is a common issue for diffusion models, and (3) weak alignment with input conditions, which is a critical issue for medical imaging. MAISI, a previously proposed framework, addresses generalizability issues but still suffers from slow inference and limited condition consistency. In this work, we present MAISI-v2, the first accelerated 3D medical image synthesis framework that integrates rectified flow to enable fast and high quality generation. To further enhance condition fidelity, we introduce a novel region-specific contrastive loss to enhance the sensitivity to region of interest. Our experiments show that MAISI-v2 can achieve SOTA image quality with $33 \times$ acceleration for latent diffusion model. We also conducted a downstream segmentation experiment to show that the synthetic images can be used for data augmentation. We release our code, training details, model weights, and a GUI demo to facilitate reproducibility and promote further development within the community.
title MAISI-v2: Accelerated 3D High-Resolution Medical Image Synthesis with Rectified Flow and Region-specific Contrastive Loss
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.05772