Vocabulary-free Image Classification and Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Conti, Alessandro, Fini, Enrico, Mancini, Massimiliano, Rota, Paolo, Wang, Yiming, Ricci, Elisa
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910412395511808
author Conti, Alessandro
Fini, Enrico
Mancini, Massimiliano
Rota, Paolo
Wang, Yiming
Ricci, Elisa
author_facet Conti, Alessandro
Fini, Enrico
Mancini, Massimiliano
Rota, Paolo
Wang, Yiming
Ricci, Elisa
contents Large vision-language models revolutionized image classification and semantic segmentation paradigms. However, they typically assume a pre-defined set of categories, or vocabulary, at test time for composing textual prompts. This assumption is impractical in scenarios with unknown or evolving semantic context. Here, we address this issue and introduce the Vocabulary-free Image Classification (VIC) task, which aims to assign a class from an unconstrained language-induced semantic space to an input image without needing a known vocabulary. VIC is challenging due to the vastness of the semantic space, which contains millions of concepts, including fine-grained categories. To address VIC, we propose Category Search from External Databases (CaSED), a training-free method that leverages a pre-trained vision-language model and an external database. CaSED first extracts the set of candidate categories from the most semantically similar captions in the database and then assigns the image to the best-matching candidate category according to the same vision-language model. Furthermore, we demonstrate that CaSED can be applied locally to generate a coarse segmentation mask that classifies image regions, introducing the task of Vocabulary-free Semantic Segmentation. CaSED and its variants outperform other more complex vision-language models, on classification and semantic segmentation benchmarks, while using much fewer parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2404_10864
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Vocabulary-free Image Classification and Semantic Segmentation
Conti, Alessandro
Fini, Enrico
Mancini, Massimiliano
Rota, Paolo
Wang, Yiming
Ricci, Elisa
Computer Vision and Pattern Recognition
Large vision-language models revolutionized image classification and semantic segmentation paradigms. However, they typically assume a pre-defined set of categories, or vocabulary, at test time for composing textual prompts. This assumption is impractical in scenarios with unknown or evolving semantic context. Here, we address this issue and introduce the Vocabulary-free Image Classification (VIC) task, which aims to assign a class from an unconstrained language-induced semantic space to an input image without needing a known vocabulary. VIC is challenging due to the vastness of the semantic space, which contains millions of concepts, including fine-grained categories. To address VIC, we propose Category Search from External Databases (CaSED), a training-free method that leverages a pre-trained vision-language model and an external database. CaSED first extracts the set of candidate categories from the most semantically similar captions in the database and then assigns the image to the best-matching candidate category according to the same vision-language model. Furthermore, we demonstrate that CaSED can be applied locally to generate a coarse segmentation mask that classifies image regions, introducing the task of Vocabulary-free Semantic Segmentation. CaSED and its variants outperform other more complex vision-language models, on classification and semantic segmentation benchmarks, while using much fewer parameters.
title Vocabulary-free Image Classification and Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.10864