ProGiDiff: Prompt-Guided Diffusion-Based Medical Image Segmentation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lin, Yuan, Xu, Murong, Hölle, Marc, Prabhakar, Chinmay, Maier, Andreas, Belagiannis, Vasileios, Menze, Bjoern, Shit, Suprosanna
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917218023899136
author Lin, Yuan
Xu, Murong
Hölle, Marc
Prabhakar, Chinmay
Maier, Andreas
Belagiannis, Vasileios
Menze, Bjoern
Shit, Suprosanna
author_facet Lin, Yuan
Xu, Murong
Hölle, Marc
Prabhakar, Chinmay
Maier, Andreas
Belagiannis, Vasileios
Menze, Bjoern
Shit, Suprosanna
contents Widely adopted medical image segmentation methods, although efficient, are primarily deterministic and remain poorly amenable to natural language prompts. Thus, they lack the capability to estimate multiple proposals, human interaction, and cross-modality adaptation. Recently, text-to-image diffusion models have shown potential to bridge the gap. However, training them from scratch requires a large dataset-a limitation for medical image segmentation. Furthermore, they are often limited to binary segmentation and cannot be conditioned on a natural language prompt. To this end, we propose a novel framework called ProGiDiff that leverages existing image generation models for medical image segmentation purposes. Specifically, we propose a ControlNet-style conditioning mechanism with a custom encoder, suitable for image conditioning, to steer a pre-trained diffusion model to output segmentation masks. It naturally extends to a multi-class setting simply by prompting the target organ. Our experiment on organ segmentation from CT images demonstrates strong performance compared to previous methods and could greatly benefit from an expert-in-the-loop setting to leverage multiple proposals. Importantly, we demonstrate that the learned conditioning mechanism can be easily transferred through low-rank, few-shot adaptation to segment MR images.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16060
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ProGiDiff: Prompt-Guided Diffusion-Based Medical Image Segmentation
Lin, Yuan
Xu, Murong
Hölle, Marc
Prabhakar, Chinmay
Maier, Andreas
Belagiannis, Vasileios
Menze, Bjoern
Shit, Suprosanna
Computer Vision and Pattern Recognition
Widely adopted medical image segmentation methods, although efficient, are primarily deterministic and remain poorly amenable to natural language prompts. Thus, they lack the capability to estimate multiple proposals, human interaction, and cross-modality adaptation. Recently, text-to-image diffusion models have shown potential to bridge the gap. However, training them from scratch requires a large dataset-a limitation for medical image segmentation. Furthermore, they are often limited to binary segmentation and cannot be conditioned on a natural language prompt. To this end, we propose a novel framework called ProGiDiff that leverages existing image generation models for medical image segmentation purposes. Specifically, we propose a ControlNet-style conditioning mechanism with a custom encoder, suitable for image conditioning, to steer a pre-trained diffusion model to output segmentation masks. It naturally extends to a multi-class setting simply by prompting the target organ. Our experiment on organ segmentation from CT images demonstrates strong performance compared to previous methods and could greatly benefit from an expert-in-the-loop setting to leverage multiple proposals. Importantly, we demonstrate that the learned conditioning mechanism can be easily transferred through low-rank, few-shot adaptation to segment MR images.
title ProGiDiff: Prompt-Guided Diffusion-Based Medical Image Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.16060