SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ye, Hanrong, Kuen, Jason, Liu, Qing, Lin, Zhe, Price, Brian, Xu, Dan
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917713211817984
author Ye, Hanrong
Kuen, Jason
Liu, Qing
Lin, Zhe
Price, Brian
Xu, Dan
author_facet Ye, Hanrong
Kuen, Jason
Liu, Qing
Lin, Zhe
Price, Brian
Xu, Dan
contents We propose SegGen, a highly-effective training data generation method for image segmentation, which pushes the performance limits of state-of-the-art segmentation models to a significant extent. SegGen designs and integrates two data generation strategies: MaskSyn and ImgSyn. (i) MaskSyn synthesizes new mask-image pairs via our proposed text-to-mask generation model and mask-to-image generation model, greatly improving the diversity in segmentation masks for model supervision; (ii) ImgSyn synthesizes new images based on existing masks using the mask-to-image generation model, strongly improving image diversity for model inputs. On the highly competitive ADE20K and COCO benchmarks, our data generation method markedly improves the performance of state-of-the-art segmentation models in semantic segmentation, panoptic segmentation, and instance segmentation. Notably, in terms of the ADE20K mIoU, Mask2Former R50 is largely boosted from 47.2 to 49.9 (+2.7); Mask2Former Swin-L is also significantly increased from 56.1 to 57.4 (+1.3). These promising results strongly suggest the effectiveness of our SegGen even when abundant human-annotated training data is utilized. Moreover, training with our synthetic data makes the segmentation models more robust towards unseen domains. Project website: https://seggenerator.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2311_03355
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis
Ye, Hanrong
Kuen, Jason
Liu, Qing
Lin, Zhe
Price, Brian
Xu, Dan
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
We propose SegGen, a highly-effective training data generation method for image segmentation, which pushes the performance limits of state-of-the-art segmentation models to a significant extent. SegGen designs and integrates two data generation strategies: MaskSyn and ImgSyn. (i) MaskSyn synthesizes new mask-image pairs via our proposed text-to-mask generation model and mask-to-image generation model, greatly improving the diversity in segmentation masks for model supervision; (ii) ImgSyn synthesizes new images based on existing masks using the mask-to-image generation model, strongly improving image diversity for model inputs. On the highly competitive ADE20K and COCO benchmarks, our data generation method markedly improves the performance of state-of-the-art segmentation models in semantic segmentation, panoptic segmentation, and instance segmentation. Notably, in terms of the ADE20K mIoU, Mask2Former R50 is largely boosted from 47.2 to 49.9 (+2.7); Mask2Former Swin-L is also significantly increased from 56.1 to 57.4 (+1.3). These promising results strongly suggest the effectiveness of our SegGen even when abundant human-annotated training data is utilized. Moreover, training with our synthetic data makes the segmentation models more robust towards unseen domains. Project website: https://seggenerator.github.io
title SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2311.03355