Latent Space Synergy: Text-Guided Data Augmentation for Direct Diffusion Biomedical Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aqeel, Muhammad, Nazir, Maham, Ruan, Zanxi, Setti, Francesco
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916853036613632
author Aqeel, Muhammad
Nazir, Maham
Ruan, Zanxi
Setti, Francesco
author_facet Aqeel, Muhammad
Nazir, Maham
Ruan, Zanxi
Setti, Francesco
contents Medical image segmentation suffers from data scarcity, particularly in polyp detection where annotation requires specialized expertise. We present SynDiff, a framework combining text-guided synthetic data generation with efficient diffusion-based segmentation. Our approach employs latent diffusion models to generate clinically realistic synthetic polyps through text-conditioned inpainting, augmenting limited training data with semantically diverse samples. Unlike traditional diffusion methods requiring iterative denoising, we introduce direct latent estimation enabling single-step inference with T x computational speedup. On CVC-ClinicDB, SynDiff achieves 96.0% Dice and 92.9% IoU while maintaining real-time capability suitable for clinical deployment. The framework demonstrates that controlled synthetic augmentation improves segmentation robustness without distribution shift. SynDiff bridges the gap between data-hungry deep learning models and clinical constraints, offering an efficient solution for deployment in resourcelimited medical settings.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15361
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Latent Space Synergy: Text-Guided Data Augmentation for Direct Diffusion Biomedical Segmentation
Aqeel, Muhammad
Nazir, Maham
Ruan, Zanxi
Setti, Francesco
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Medical image segmentation suffers from data scarcity, particularly in polyp detection where annotation requires specialized expertise. We present SynDiff, a framework combining text-guided synthetic data generation with efficient diffusion-based segmentation. Our approach employs latent diffusion models to generate clinically realistic synthetic polyps through text-conditioned inpainting, augmenting limited training data with semantically diverse samples. Unlike traditional diffusion methods requiring iterative denoising, we introduce direct latent estimation enabling single-step inference with T x computational speedup. On CVC-ClinicDB, SynDiff achieves 96.0% Dice and 92.9% IoU while maintaining real-time capability suitable for clinical deployment. The framework demonstrates that controlled synthetic augmentation improves segmentation robustness without distribution shift. SynDiff bridges the gap between data-hungry deep learning models and clinical constraints, offering an efficient solution for deployment in resourcelimited medical settings.
title Latent Space Synergy: Text-Guided Data Augmentation for Direct Diffusion Biomedical Segmentation
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.15361