VAP-Diffusion: Enriching Descriptions with MLLMs for Enhanced Medical Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Peng, Fu, Junhu, Guo, Bowen, Li, Zeju, Wang, Yuanyuan, Guo, Yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916817106108416
author Huang, Peng
Fu, Junhu
Guo, Bowen
Li, Zeju
Wang, Yuanyuan
Guo, Yi
author_facet Huang, Peng
Fu, Junhu
Guo, Bowen
Li, Zeju
Wang, Yuanyuan
Guo, Yi
contents As the appearance of medical images is influenced by multiple underlying factors, generative models require rich attribute information beyond labels to produce realistic and diverse images. For instance, generating an image of skin lesion with specific patterns demands descriptions that go beyond diagnosis, such as shape, size, texture, and color. However, such detailed descriptions are not always accessible. To address this, we explore a framework, termed Visual Attribute Prompts (VAP)-Diffusion, to leverage external knowledge from pre-trained Multi-modal Large Language Models (MLLMs) to improve the quality and diversity of medical image generation. First, to derive descriptions from MLLMs without hallucination, we design a series of prompts following Chain-of-Thoughts for common medical imaging tasks, including dermatologic, colorectal, and chest X-ray images. Generated descriptions are utilized during training and stored across different categories. During testing, descriptions are randomly retrieved from the corresponding category for inference. Moreover, to make the generator robust to unseen combination of descriptions at the test time, we propose a Prototype Condition Mechanism that restricts test embeddings to be similar to those from training. Experiments on three common types of medical imaging across four datasets verify the effectiveness of VAP-Diffusion.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23641
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VAP-Diffusion: Enriching Descriptions with MLLMs for Enhanced Medical Image Generation
Huang, Peng
Fu, Junhu
Guo, Bowen
Li, Zeju
Wang, Yuanyuan
Guo, Yi
Computer Vision and Pattern Recognition
Artificial Intelligence
As the appearance of medical images is influenced by multiple underlying factors, generative models require rich attribute information beyond labels to produce realistic and diverse images. For instance, generating an image of skin lesion with specific patterns demands descriptions that go beyond diagnosis, such as shape, size, texture, and color. However, such detailed descriptions are not always accessible. To address this, we explore a framework, termed Visual Attribute Prompts (VAP)-Diffusion, to leverage external knowledge from pre-trained Multi-modal Large Language Models (MLLMs) to improve the quality and diversity of medical image generation. First, to derive descriptions from MLLMs without hallucination, we design a series of prompts following Chain-of-Thoughts for common medical imaging tasks, including dermatologic, colorectal, and chest X-ray images. Generated descriptions are utilized during training and stored across different categories. During testing, descriptions are randomly retrieved from the corresponding category for inference. Moreover, to make the generator robust to unseen combination of descriptions at the test time, we propose a Prototype Condition Mechanism that restricts test embeddings to be similar to those from training. Experiments on three common types of medical imaging across four datasets verify the effectiveness of VAP-Diffusion.
title VAP-Diffusion: Enriching Descriptions with MLLMs for Enhanced Medical Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.23641