From Open-Vocabulary to Vocabulary-Free Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Reichard, Klara, Rizzoli, Giulia, Gasperini, Stefano, Hoyer, Lukas, Zanuttigh, Pietro, Navab, Nassir, Tombari, Federico
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913693854334976
author Reichard, Klara
Rizzoli, Giulia
Gasperini, Stefano
Hoyer, Lukas
Zanuttigh, Pietro
Navab, Nassir
Tombari, Federico
author_facet Reichard, Klara
Rizzoli, Giulia
Gasperini, Stefano
Hoyer, Lukas
Zanuttigh, Pietro
Navab, Nassir
Tombari, Federico
contents Open-vocabulary semantic segmentation enables models to identify novel object categories beyond their training data. While this flexibility represents a significant advancement, current approaches still rely on manually specified class names as input, creating an inherent bottleneck in real-world applications. This work proposes a Vocabulary-Free Semantic Segmentation pipeline, eliminating the need for predefined class vocabularies. Specifically, we address the chicken-and-egg problem where users need knowledge of all potential objects within a scene to identify them, yet the purpose of segmentation is often to discover these objects. The proposed approach leverages Vision-Language Models to automatically recognize objects and generate appropriate class names, aiming to solve the challenge of class specification and naming quality. Through extensive experiments on several public datasets, we highlight the crucial role of the text encoder in model performance, particularly when the image text classes are paired with generated descriptions. Despite the challenges introduced by the sensitivity of the segmentation text encoder to false negatives within the class tagging process, which adds complexity to the task, we demonstrate that our fully automated pipeline significantly enhances vocabulary-free segmentation accuracy across diverse real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11891
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Open-Vocabulary to Vocabulary-Free Semantic Segmentation
Reichard, Klara
Rizzoli, Giulia
Gasperini, Stefano
Hoyer, Lukas
Zanuttigh, Pietro
Navab, Nassir
Tombari, Federico
Computer Vision and Pattern Recognition
Open-vocabulary semantic segmentation enables models to identify novel object categories beyond their training data. While this flexibility represents a significant advancement, current approaches still rely on manually specified class names as input, creating an inherent bottleneck in real-world applications. This work proposes a Vocabulary-Free Semantic Segmentation pipeline, eliminating the need for predefined class vocabularies. Specifically, we address the chicken-and-egg problem where users need knowledge of all potential objects within a scene to identify them, yet the purpose of segmentation is often to discover these objects. The proposed approach leverages Vision-Language Models to automatically recognize objects and generate appropriate class names, aiming to solve the challenge of class specification and naming quality. Through extensive experiments on several public datasets, we highlight the crucial role of the text encoder in model performance, particularly when the image text classes are paired with generated descriptions. Despite the challenges introduced by the sensitivity of the segmentation text encoder to false negatives within the class tagging process, which adds complexity to the task, we demonstrate that our fully automated pipeline significantly enhances vocabulary-free segmentation accuracy across diverse real-world scenarios.
title From Open-Vocabulary to Vocabulary-Free Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.11891