PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask Conditions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bhosale, Mahesh, Wasi, Abdul, Zhai, Yuanhao, Tian, Yunjie, Border, Samuel, Xi, Nan, Sarder, Pinaki, Yuan, Junsong, Doermann, David, Gong, Xuan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913918950047744
author Bhosale, Mahesh
Wasi, Abdul
Zhai, Yuanhao
Tian, Yunjie
Border, Samuel
Xi, Nan
Sarder, Pinaki
Yuan, Junsong
Doermann, David
Gong, Xuan
author_facet Bhosale, Mahesh
Wasi, Abdul
Zhai, Yuanhao
Tian, Yunjie
Border, Samuel
Xi, Nan
Sarder, Pinaki
Yuan, Junsong
Doermann, David
Gong, Xuan
contents Diffusion-based generative models have shown promise in synthesizing histopathology images to address data scarcity caused by privacy constraints. Diagnostic text reports provide high-level semantic descriptions, and masks offer fine-grained spatial structures essential for representing distinct morphological regions. However, public datasets lack paired text and mask data for the same histopathological images, limiting their joint use in image generation. This constraint restricts the ability to fully exploit the benefits of combining both modalities for enhanced control over semantics and spatial details. To overcome this, we propose PathDiff, a diffusion framework that effectively learns from unpaired mask-text data by integrating both modalities into a unified conditioning space. PathDiff allows precise control over structural and contextual features, generating high-quality, semantically accurate images. PathDiff also improves image fidelity, text-image alignment, and faithfulness, enhancing data augmentation for downstream tasks like nuclei segmentation and classification. Extensive experiments demonstrate its superiority over existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23440
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask Conditions
Bhosale, Mahesh
Wasi, Abdul
Zhai, Yuanhao
Tian, Yunjie
Border, Samuel
Xi, Nan
Sarder, Pinaki
Yuan, Junsong
Doermann, David
Gong, Xuan
Computer Vision and Pattern Recognition
Diffusion-based generative models have shown promise in synthesizing histopathology images to address data scarcity caused by privacy constraints. Diagnostic text reports provide high-level semantic descriptions, and masks offer fine-grained spatial structures essential for representing distinct morphological regions. However, public datasets lack paired text and mask data for the same histopathological images, limiting their joint use in image generation. This constraint restricts the ability to fully exploit the benefits of combining both modalities for enhanced control over semantics and spatial details. To overcome this, we propose PathDiff, a diffusion framework that effectively learns from unpaired mask-text data by integrating both modalities into a unified conditioning space. PathDiff allows precise control over structural and contextual features, generating high-quality, semantically accurate images. PathDiff also improves image fidelity, text-image alignment, and faithfulness, enhancing data augmentation for downstream tasks like nuclei segmentation and classification. Extensive experiments demonstrate its superiority over existing methods.
title PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask Conditions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.23440