Med-Art: Diffusion Transformer for 2D Medical Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Changlu, Christensen, Anders Nymark, Hannemose, Morten Rieger
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909659967782912
author Guo, Changlu
Christensen, Anders Nymark
Hannemose, Morten Rieger
author_facet Guo, Changlu
Christensen, Anders Nymark
Hannemose, Morten Rieger
contents Text-to-image generative models have achieved remarkable breakthroughs in recent years. However, their application in medical image generation still faces significant challenges, including small dataset sizes, and scarcity of medical textual data. To address these challenges, we propose Med-Art, a framework specifically designed for medical image generation with limited data. Med-Art leverages vision-language models to generate visual descriptions of medical images which overcomes the scarcity of applicable medical textual data. Med-Art adapts a large-scale pre-trained text-to-image model, PixArt-$α$, based on the Diffusion Transformer (DiT), achieving high performance under limited data. Furthermore, we propose an innovative Hybrid-Level Diffusion Fine-tuning (HLDF) method, which enables pixel-level losses, effectively addressing issues such as overly saturated colors. We achieve state-of-the-art performance on two medical image datasets, measured by FID, KID, and downstream classification performance.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20449
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Med-Art: Diffusion Transformer for 2D Medical Text-to-Image Generation
Guo, Changlu
Christensen, Anders Nymark
Hannemose, Morten Rieger
Computer Vision and Pattern Recognition
Text-to-image generative models have achieved remarkable breakthroughs in recent years. However, their application in medical image generation still faces significant challenges, including small dataset sizes, and scarcity of medical textual data. To address these challenges, we propose Med-Art, a framework specifically designed for medical image generation with limited data. Med-Art leverages vision-language models to generate visual descriptions of medical images which overcomes the scarcity of applicable medical textual data. Med-Art adapts a large-scale pre-trained text-to-image model, PixArt-$α$, based on the Diffusion Transformer (DiT), achieving high performance under limited data. Furthermore, we propose an innovative Hybrid-Level Diffusion Fine-tuning (HLDF) method, which enables pixel-level losses, effectively addressing issues such as overly saturated colors. We achieve state-of-the-art performance on two medical image datasets, measured by FID, KID, and downstream classification performance.
title Med-Art: Diffusion Transformer for 2D Medical Text-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.20449