Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yong, Wu, SongLi, Bai, Sule, Wang, Jiahao, Wang, Yitong, Tang, Yansong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918068883554304
author Liu, Yong
Wu, SongLi
Bai, Sule
Wang, Jiahao
Wang, Yitong
Tang, Yansong
author_facet Liu, Yong
Wu, SongLi
Bai, Sule
Wang, Jiahao
Wang, Yitong
Tang, Yansong
contents Open-vocabulary segmentation aims to achieve segmentation of arbitrary categories given unlimited text inputs as guidance. To achieve this, recent works have focused on developing various technical routes to exploit the potential of large-scale pre-trained vision-language models and have made significant progress on existing benchmarks. However, we find that existing test sets are limited in measuring the models' comprehension of ``open-vocabulary" concepts, as their semantic space closely resembles the training space, even with many overlapping categories. To this end, we present a new benchmark named OpenBench that differs significantly from the training semantics. It is designed to better assess the model's ability to understand and segment a wide range of real-world concepts. When testing existing methods on OpenBench, we find that their performance diverges from the conclusions drawn on existing test sets. In addition, we propose a method named OVSNet to improve the segmentation performance for diverse and open scenarios. Through elaborate fusion of heterogeneous features and cost-free expansion of the training space, OVSNet achieves state-of-the-art results on both existing datasets and our proposed OpenBench. Corresponding analysis demonstrate the soundness and effectiveness of our proposed benchmark and method.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16058
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation
Liu, Yong
Wu, SongLi
Bai, Sule
Wang, Jiahao
Wang, Yitong
Tang, Yansong
Computer Vision and Pattern Recognition
Open-vocabulary segmentation aims to achieve segmentation of arbitrary categories given unlimited text inputs as guidance. To achieve this, recent works have focused on developing various technical routes to exploit the potential of large-scale pre-trained vision-language models and have made significant progress on existing benchmarks. However, we find that existing test sets are limited in measuring the models' comprehension of ``open-vocabulary" concepts, as their semantic space closely resembles the training space, even with many overlapping categories. To this end, we present a new benchmark named OpenBench that differs significantly from the training semantics. It is designed to better assess the model's ability to understand and segment a wide range of real-world concepts. When testing existing methods on OpenBench, we find that their performance diverges from the conclusions drawn on existing test sets. In addition, we propose a method named OVSNet to improve the segmentation performance for diverse and open scenarios. Through elaborate fusion of heterogeneous features and cost-free expansion of the training space, OVSNet achieves state-of-the-art results on both existing datasets and our proposed OpenBench. Corresponding analysis demonstrate the soundness and effectiveness of our proposed benchmark and method.
title Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.16058