CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cho, Seokju, Shin, Heeseong, Hong, Sunghwan, Arnab, Anurag, Seo, Paul Hongsuck, Kim, Seungryong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917626350927872
author Cho, Seokju
Shin, Heeseong
Hong, Sunghwan
Arnab, Anurag
Seo, Paul Hongsuck
Kim, Seungryong
author_facet Cho, Seokju
Shin, Heeseong
Hong, Sunghwan
Arnab, Anurag
Seo, Paul Hongsuck
Kim, Seungryong
contents Open-vocabulary semantic segmentation presents the challenge of labeling each pixel within an image based on a wide range of text descriptions. In this work, we introduce a novel cost-based approach to adapt vision-language foundation models, notably CLIP, for the intricate task of semantic segmentation. Through aggregating the cosine similarity score, i.e., the cost volume between image and text embeddings, our method potently adapts CLIP for segmenting seen and unseen classes by fine-tuning its encoders, addressing the challenges faced by existing methods in handling unseen classes. Building upon this, we explore methods to effectively aggregate the cost volume considering its multi-modal nature of being established between image and text embeddings. Furthermore, we examine various methods for efficiently fine-tuning CLIP.
format Preprint
id arxiv_https___arxiv_org_abs_2303_11797
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation
Cho, Seokju
Shin, Heeseong
Hong, Sunghwan
Arnab, Anurag
Seo, Paul Hongsuck
Kim, Seungryong
Computer Vision and Pattern Recognition
Open-vocabulary semantic segmentation presents the challenge of labeling each pixel within an image based on a wide range of text descriptions. In this work, we introduce a novel cost-based approach to adapt vision-language foundation models, notably CLIP, for the intricate task of semantic segmentation. Through aggregating the cosine similarity score, i.e., the cost volume between image and text embeddings, our method potently adapts CLIP for segmenting seen and unseen classes by fine-tuning its encoders, addressing the challenges faced by existing methods in handling unseen classes. Building upon this, we explore methods to effectively aggregate the cost volume considering its multi-modal nature of being established between image and text embeddings. Furthermore, we examine various methods for efficiently fine-tuning CLIP.
title CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2303.11797