Multi-Modal Prototypes for Open-World Semantic Segmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Yuhuan, Ma, Chaofan, Ju, Chen, Zhang, Fei, Yao, Jiangchao, Zhang, Ya, Wang, Yanfeng
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916319350226944
author Yang, Yuhuan
Ma, Chaofan
Ju, Chen
Zhang, Fei
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
author_facet Yang, Yuhuan
Ma, Chaofan
Ju, Chen
Zhang, Fei
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
contents In semantic segmentation, generalizing a visual system to both seen categories and novel categories at inference time has always been practically valuable yet challenging. To enable such functionality, existing methods mainly rely on either providing several support demonstrations from the visual aspect or characterizing the informative clues from the textual aspect (e.g., the class names). Nevertheless, both two lines neglect the complementary intrinsic of low-level visual and high-level language information, while the explorations that consider visual and textual modalities as a whole to promote predictions are still limited. To close this gap, we propose to encompass textual and visual clues as multi-modal prototypes to allow more comprehensive support for open-world semantic segmentation, and build a novel prototype-based segmentation framework to realize this promise. To be specific, unlike the straightforward combination of bi-modal clues, we decompose the high-level language information as multi-aspect prototypes and aggregate the low-level visual information as more semantic prototypes, on basis of which, a fine-grained complementary fusion makes the multi-modal prototypes more powerful and accurate to promote the prediction. Based on an elastic mask prediction module that permits any number and form of prototype inputs, we are able to solve the zero-shot, few-shot and generalized counterpart tasks in one architecture. Extensive experiments on both PASCAL-$5^i$ and COCO-$20^i$ datasets show the consistent superiority of the proposed method compared with the previous state-of-the-art approaches, and a range of ablation studies thoroughly dissects each component in our framework both quantitatively and qualitatively that verify their effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2307_02003
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Multi-Modal Prototypes for Open-World Semantic Segmentation
Yang, Yuhuan
Ma, Chaofan
Ju, Chen
Zhang, Fei
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
Computer Vision and Pattern Recognition
In semantic segmentation, generalizing a visual system to both seen categories and novel categories at inference time has always been practically valuable yet challenging. To enable such functionality, existing methods mainly rely on either providing several support demonstrations from the visual aspect or characterizing the informative clues from the textual aspect (e.g., the class names). Nevertheless, both two lines neglect the complementary intrinsic of low-level visual and high-level language information, while the explorations that consider visual and textual modalities as a whole to promote predictions are still limited. To close this gap, we propose to encompass textual and visual clues as multi-modal prototypes to allow more comprehensive support for open-world semantic segmentation, and build a novel prototype-based segmentation framework to realize this promise. To be specific, unlike the straightforward combination of bi-modal clues, we decompose the high-level language information as multi-aspect prototypes and aggregate the low-level visual information as more semantic prototypes, on basis of which, a fine-grained complementary fusion makes the multi-modal prototypes more powerful and accurate to promote the prediction. Based on an elastic mask prediction module that permits any number and form of prototype inputs, we are able to solve the zero-shot, few-shot and generalized counterpart tasks in one architecture. Extensive experiments on both PASCAL-$5^i$ and COCO-$20^i$ datasets show the consistent superiority of the proposed method compared with the previous state-of-the-art approaches, and a range of ablation studies thoroughly dissects each component in our framework both quantitatively and qualitatively that verify their effectiveness.
title Multi-Modal Prototypes for Open-World Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2307.02003