Vision-Language Synthetic Data Enhances Echocardiography Downstream Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ashrafian, Pooria, Yazdani, Milad, Heidari, Moein, Shahriari, Dena, Hacihaliloglu, Ilker
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911818896637952
author Ashrafian, Pooria
Yazdani, Milad
Heidari, Moein
Shahriari, Dena
Hacihaliloglu, Ilker
author_facet Ashrafian, Pooria
Yazdani, Milad
Heidari, Moein
Shahriari, Dena
Hacihaliloglu, Ilker
contents High-quality, large-scale data is essential for robust deep learning models in medical applications, particularly ultrasound image analysis. Diffusion models facilitate high-fidelity medical image generation, reducing the costs associated with acquiring and annotating new images. This paper utilizes recent vision-language models to produce diverse and realistic synthetic echocardiography image data, preserving key features of the original images guided by textual and semantic label maps. Specifically, we investigate three potential avenues: unconditional generation, generation guided by text, and a hybrid approach incorporating both textual and semantic supervision. We show that the rich contextual information present in the synthesized data potentially enhances the accuracy and interpretability of downstream tasks, such as echocardiography segmentation and classification with improved metrics and faster convergence. Our implementation with checkpoints, prompts, and the created synthetic dataset will be publicly available at \href{https://github.com/Pooria90/DiffEcho}{GitHub}.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19880
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Vision-Language Synthetic Data Enhances Echocardiography Downstream Tasks
Ashrafian, Pooria
Yazdani, Milad
Heidari, Moein
Shahriari, Dena
Hacihaliloglu, Ilker
Image and Video Processing
Computer Vision and Pattern Recognition
High-quality, large-scale data is essential for robust deep learning models in medical applications, particularly ultrasound image analysis. Diffusion models facilitate high-fidelity medical image generation, reducing the costs associated with acquiring and annotating new images. This paper utilizes recent vision-language models to produce diverse and realistic synthetic echocardiography image data, preserving key features of the original images guided by textual and semantic label maps. Specifically, we investigate three potential avenues: unconditional generation, generation guided by text, and a hybrid approach incorporating both textual and semantic supervision. We show that the rich contextual information present in the synthesized data potentially enhances the accuracy and interpretability of downstream tasks, such as echocardiography segmentation and classification with improved metrics and faster convergence. Our implementation with checkpoints, prompts, and the created synthetic dataset will be publicly available at \href{https://github.com/Pooria90/DiffEcho}{GitHub}.
title Vision-Language Synthetic Data Enhances Echocardiography Downstream Tasks
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.19880