Exploring Limits of Diffusion-Synthetic Training with Weakly Supervised Semantic Segmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yoshihashi, Ryota, Otsuka, Yuya, Doi, Kenji, Tanaka, Tomohiro, Kataoka, Hirokatsu
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911839887032320
author Yoshihashi, Ryota
Otsuka, Yuya
Doi, Kenji
Tanaka, Tomohiro
Kataoka, Hirokatsu
author_facet Yoshihashi, Ryota
Otsuka, Yuya
Doi, Kenji
Tanaka, Tomohiro
Kataoka, Hirokatsu
contents The advance of generative models for images has inspired various training techniques for image recognition utilizing synthetic images. In semantic segmentation, one promising approach is extracting pseudo-masks from attention maps in text-to-image diffusion models, which enables real-image-and-annotation-free training. However, the pioneering training method using the diffusion-synthetic images and pseudo-masks, i.e., DiffuMask has limitations in terms of mask quality, scalability, and ranges of applicable domains. To overcome these limitations, this work introduces three techniques for diffusion-synthetic semantic segmentation training. First, reliability-aware robust training, originally used in weakly supervised learning, helps segmentation with insufficient synthetic mask quality. %Second, large-scale pretraining of whole segmentation models, not only backbones, on synthetic ImageNet-1k-class images with pixel-labels benefits downstream segmentation tasks. Second, we introduce prompt augmentation, data augmentation to the prompt text set to scale up and diversify training images with a limited text resources. Finally, LoRA-based adaptation of Stable Diffusion enables the transfer to a distant domain, e.g., auto-driving images. Experiments in PASCAL VOC, ImageNet-S, and Cityscapes show that our method effectively closes gap between real and synthetic training in semantic segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2309_01369
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Exploring Limits of Diffusion-Synthetic Training with Weakly Supervised Semantic Segmentation
Yoshihashi, Ryota
Otsuka, Yuya
Doi, Kenji
Tanaka, Tomohiro
Kataoka, Hirokatsu
Computer Vision and Pattern Recognition
The advance of generative models for images has inspired various training techniques for image recognition utilizing synthetic images. In semantic segmentation, one promising approach is extracting pseudo-masks from attention maps in text-to-image diffusion models, which enables real-image-and-annotation-free training. However, the pioneering training method using the diffusion-synthetic images and pseudo-masks, i.e., DiffuMask has limitations in terms of mask quality, scalability, and ranges of applicable domains. To overcome these limitations, this work introduces three techniques for diffusion-synthetic semantic segmentation training. First, reliability-aware robust training, originally used in weakly supervised learning, helps segmentation with insufficient synthetic mask quality. %Second, large-scale pretraining of whole segmentation models, not only backbones, on synthetic ImageNet-1k-class images with pixel-labels benefits downstream segmentation tasks. Second, we introduce prompt augmentation, data augmentation to the prompt text set to scale up and diversify training images with a limited text resources. Finally, LoRA-based adaptation of Stable Diffusion enables the transfer to a distant domain, e.g., auto-driving images. Experiments in PASCAL VOC, ImageNet-S, and Cityscapes show that our method effectively closes gap between real and synthetic training in semantic segmentation.
title Exploring Limits of Diffusion-Synthetic Training with Weakly Supervised Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2309.01369