Mask Factory: Towards High-quality Synthetic Data Generation for Dichotomous Image Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qian, Haotian, Chen, YD, Lou, Shengtao, Khan, Fahad Shahbaz, Jin, Xiaogang, Fan, Deng-Ping
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910764495798272
author Qian, Haotian
Chen, YD
Lou, Shengtao
Khan, Fahad Shahbaz
Jin, Xiaogang
Fan, Deng-Ping
author_facet Qian, Haotian
Chen, YD
Lou, Shengtao
Khan, Fahad Shahbaz
Jin, Xiaogang
Fan, Deng-Ping
contents Dichotomous Image Segmentation (DIS) tasks require highly precise annotations, and traditional dataset creation methods are labor intensive, costly, and require extensive domain expertise. Although using synthetic data for DIS is a promising solution to these challenges, current generative models and techniques struggle with the issues of scene deviations, noise-induced errors, and limited training sample variability. To address these issues, we introduce a novel approach, \textbf{\ourmodel{}}, which provides a scalable solution for generating diverse and precise datasets, markedly reducing preparation time and costs. We first introduce a general mask editing method that combines rigid and non-rigid editing techniques to generate high-quality synthetic masks. Specially, rigid editing leverages geometric priors from diffusion models to achieve precise viewpoint transformations under zero-shot conditions, while non-rigid editing employs adversarial training and self-attention mechanisms for complex, topologically consistent modifications. Then, we generate pairs of high-resolution image and accurate segmentation mask using a multi-conditional control generation method. Finally, our experiments on the widely-used DIS5K dataset benchmark demonstrate superior performance in quality and efficiency compared to existing methods. The code is available at \url{https://qian-hao-tian.github.io/MaskFactory/}.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19080
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mask Factory: Towards High-quality Synthetic Data Generation for Dichotomous Image Segmentation
Qian, Haotian
Chen, YD
Lou, Shengtao
Khan, Fahad Shahbaz
Jin, Xiaogang
Fan, Deng-Ping
Computer Vision and Pattern Recognition
Dichotomous Image Segmentation (DIS) tasks require highly precise annotations, and traditional dataset creation methods are labor intensive, costly, and require extensive domain expertise. Although using synthetic data for DIS is a promising solution to these challenges, current generative models and techniques struggle with the issues of scene deviations, noise-induced errors, and limited training sample variability. To address these issues, we introduce a novel approach, \textbf{\ourmodel{}}, which provides a scalable solution for generating diverse and precise datasets, markedly reducing preparation time and costs. We first introduce a general mask editing method that combines rigid and non-rigid editing techniques to generate high-quality synthetic masks. Specially, rigid editing leverages geometric priors from diffusion models to achieve precise viewpoint transformations under zero-shot conditions, while non-rigid editing employs adversarial training and self-attention mechanisms for complex, topologically consistent modifications. Then, we generate pairs of high-resolution image and accurate segmentation mask using a multi-conditional control generation method. Finally, our experiments on the widely-used DIS5K dataset benchmark demonstrate superior performance in quality and efficiency compared to existing methods. The code is available at \url{https://qian-hao-tian.github.io/MaskFactory/}.
title Mask Factory: Towards High-quality Synthetic Data Generation for Dichotomous Image Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.19080