A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xiao, Junfei, Zhou, Ziqi, Li, Wenxuan, Lan, Shiyi, Mei, Jieru, Yu, Zhiding, Yuille, Alan, Zhou, Yuyin, Xie, Cihang
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916357727059968
author Xiao, Junfei
Zhou, Ziqi
Li, Wenxuan
Lan, Shiyi
Mei, Jieru
Yu, Zhiding
Yuille, Alan
Zhou, Yuyin
Xie, Cihang
author_facet Xiao, Junfei
Zhou, Ziqi
Li, Wenxuan
Lan, Shiyi
Mei, Jieru
Yu, Zhiding
Yuille, Alan
Zhou, Yuyin
Xie, Cihang
contents This paper introduces ProLab, a novel approach using property-level label space for creating strong interpretable segmentation models. Instead of relying solely on category-specific annotations, ProLab uses descriptive properties grounded in common sense knowledge for supervising segmentation models. It is based on two core designs. First, we employ Large Language Models (LLMs) and carefully crafted prompts to generate descriptions of all involved categories that carry meaningful common sense knowledge and follow a structured format. Second, we introduce a description embedding model preserving semantic correlation across descriptions and then cluster them into a set of descriptive properties (e.g., 256) using K-Means. These properties are based on interpretable common sense knowledge consistent with theories of human recognition. We empirically show that our approach makes segmentation models perform stronger on five classic benchmarks (e.g., ADE20K, COCO-Stuff, Pascal Context, Cityscapes, and BDD). Our method also shows better scalability with extended training steps than category-level supervision. Our interpretable segmentation framework also emerges with the generalization ability to segment out-of-domain or unknown categories using only in-domain descriptive properties. Code is available at https://github.com/lambert-x/ProLab.
format Preprint
id arxiv_https___arxiv_org_abs_2312_13764
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties
Xiao, Junfei
Zhou, Ziqi
Li, Wenxuan
Lan, Shiyi
Mei, Jieru
Yu, Zhiding
Yuille, Alan
Zhou, Yuyin
Xie, Cihang
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
This paper introduces ProLab, a novel approach using property-level label space for creating strong interpretable segmentation models. Instead of relying solely on category-specific annotations, ProLab uses descriptive properties grounded in common sense knowledge for supervising segmentation models. It is based on two core designs. First, we employ Large Language Models (LLMs) and carefully crafted prompts to generate descriptions of all involved categories that carry meaningful common sense knowledge and follow a structured format. Second, we introduce a description embedding model preserving semantic correlation across descriptions and then cluster them into a set of descriptive properties (e.g., 256) using K-Means. These properties are based on interpretable common sense knowledge consistent with theories of human recognition. We empirically show that our approach makes segmentation models perform stronger on five classic benchmarks (e.g., ADE20K, COCO-Stuff, Pascal Context, Cityscapes, and BDD). Our method also shows better scalability with extended training steps than category-level supervision. Our interpretable segmentation framework also emerges with the generalization ability to segment out-of-domain or unknown categories using only in-domain descriptive properties. Code is available at https://github.com/lambert-x/ProLab.
title A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2312.13764