MaskDiffusion: Exploiting Pre-trained Diffusion Models for Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kawano, Yasufumi, Aoki, Yoshimitsu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910370557329408
author Kawano, Yasufumi
Aoki, Yoshimitsu
author_facet Kawano, Yasufumi
Aoki, Yoshimitsu
contents Semantic segmentation is essential in computer vision for various applications, yet traditional approaches face significant challenges, including the high cost of annotation and extensive training for supervised learning. Additionally, due to the limited predefined categories in supervised learning, models typically struggle with infrequent classes and are unable to predict novel classes. To address these limitations, we propose MaskDiffusion, an innovative approach that leverages pretrained frozen Stable Diffusion to achieve open-vocabulary semantic segmentation without the need for additional training or annotation, leading to improved performance compared to similar methods. We also demonstrate the superior performance of MaskDiffusion in handling open vocabularies, including fine-grained and proper noun-based categories, thus expanding the scope of segmentation applications. Overall, our MaskDiffusion shows significant qualitative and quantitative improvements in contrast to other comparable unsupervised segmentation methods, i.e. on the Potsdam dataset (+10.5 mIoU compared to GEM) and COCO-Stuff (+14.8 mIoU compared to DiffSeg). All code and data will be released at https://github.com/Valkyrja3607/MaskDiffusion.
format Preprint
id arxiv_https___arxiv_org_abs_2403_11194
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MaskDiffusion: Exploiting Pre-trained Diffusion Models for Semantic Segmentation
Kawano, Yasufumi
Aoki, Yoshimitsu
Computer Vision and Pattern Recognition
Semantic segmentation is essential in computer vision for various applications, yet traditional approaches face significant challenges, including the high cost of annotation and extensive training for supervised learning. Additionally, due to the limited predefined categories in supervised learning, models typically struggle with infrequent classes and are unable to predict novel classes. To address these limitations, we propose MaskDiffusion, an innovative approach that leverages pretrained frozen Stable Diffusion to achieve open-vocabulary semantic segmentation without the need for additional training or annotation, leading to improved performance compared to similar methods. We also demonstrate the superior performance of MaskDiffusion in handling open vocabularies, including fine-grained and proper noun-based categories, thus expanding the scope of segmentation applications. Overall, our MaskDiffusion shows significant qualitative and quantitative improvements in contrast to other comparable unsupervised segmentation methods, i.e. on the Potsdam dataset (+10.5 mIoU compared to GEM) and COCO-Stuff (+14.8 mIoU compared to DiffSeg). All code and data will be released at https://github.com/Valkyrja3607/MaskDiffusion.
title MaskDiffusion: Exploiting Pre-trained Diffusion Models for Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.11194