ReME: A Data-Centric Framework for Training-Free Open-Vocabulary Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xuan, Xiwei, Deng, Ziquan, Ma, Kwan-Liu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908424968601600
author Xuan, Xiwei
Deng, Ziquan
Ma, Kwan-Liu
author_facet Xuan, Xiwei
Deng, Ziquan
Ma, Kwan-Liu
contents Training-free open-vocabulary semantic segmentation (OVS) aims to segment images given a set of arbitrary textual categories without costly model fine-tuning. Existing solutions often explore attention mechanisms of pre-trained models, such as CLIP, or generate synthetic data and design complex retrieval processes to perform OVS. However, their performance is limited by the capability of reliant models or the suboptimal quality of reference sets. In this work, we investigate the largely overlooked data quality problem for this challenging dense scene understanding task, and identify that a high-quality reference set can significantly benefit training-free OVS. With this observation, we introduce a data-quality-oriented framework, comprising a data pipeline to construct a reference set with well-paired segment-text embeddings and a simple similarity-based retrieval to unveil the essential effect of data. Remarkably, extensive evaluations on ten benchmark datasets demonstrate that our method outperforms all existing training-free OVS approaches, highlighting the importance of data-centric design for advancing OVS without training. Our code is available at https://github.com/xiweix/ReME .
format Preprint
id arxiv_https___arxiv_org_abs_2506_21233
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReME: A Data-Centric Framework for Training-Free Open-Vocabulary Segmentation
Xuan, Xiwei
Deng, Ziquan
Ma, Kwan-Liu
Computer Vision and Pattern Recognition
Training-free open-vocabulary semantic segmentation (OVS) aims to segment images given a set of arbitrary textual categories without costly model fine-tuning. Existing solutions often explore attention mechanisms of pre-trained models, such as CLIP, or generate synthetic data and design complex retrieval processes to perform OVS. However, their performance is limited by the capability of reliant models or the suboptimal quality of reference sets. In this work, we investigate the largely overlooked data quality problem for this challenging dense scene understanding task, and identify that a high-quality reference set can significantly benefit training-free OVS. With this observation, we introduce a data-quality-oriented framework, comprising a data pipeline to construct a reference set with well-paired segment-text embeddings and a simple similarity-based retrieval to unveil the essential effect of data. Remarkably, extensive evaluations on ten benchmark datasets demonstrate that our method outperforms all existing training-free OVS approaches, highlighting the importance of data-centric design for advancing OVS without training. Our code is available at https://github.com/xiweix/ReME .
title ReME: A Data-Centric Framework for Training-Free Open-Vocabulary Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.21233