Extending CLIP's Image-Text Alignment to Referring Image Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Seoyeon, Kang, Minguk, Kim, Dongwon, Park, Jaesik, Kwak, Suha
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911829238743040
author Kim, Seoyeon
Kang, Minguk
Kim, Dongwon
Park, Jaesik
Kwak, Suha
author_facet Kim, Seoyeon
Kang, Minguk
Kim, Dongwon
Park, Jaesik
Kwak, Suha
contents Referring Image Segmentation (RIS) is a cross-modal task that aims to segment an instance described by a natural language expression. Recent methods leverage large-scale pretrained unimodal models as backbones along with fusion techniques for joint reasoning across modalities. However, the inherent cross-modal nature of RIS raises questions about the effectiveness of unimodal backbones. We propose RISCLIP, a novel framework that effectively leverages the cross-modal nature of CLIP for RIS. Observing CLIP's inherent alignment between image and text features, we capitalize on this starting point and introduce simple but strong modules that enhance unimodal feature extraction and leverage rich alignment knowledge in CLIP's image-text shared-embedding space. RISCLIP exhibits outstanding results on all three major RIS benchmarks and also outperforms previous CLIP-based methods, demonstrating the efficacy of our strategy in extending CLIP's image-text alignment to RIS.
format Preprint
id arxiv_https___arxiv_org_abs_2306_08498
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Extending CLIP's Image-Text Alignment to Referring Image Segmentation
Kim, Seoyeon
Kang, Minguk
Kim, Dongwon
Park, Jaesik
Kwak, Suha
Computer Vision and Pattern Recognition
Referring Image Segmentation (RIS) is a cross-modal task that aims to segment an instance described by a natural language expression. Recent methods leverage large-scale pretrained unimodal models as backbones along with fusion techniques for joint reasoning across modalities. However, the inherent cross-modal nature of RIS raises questions about the effectiveness of unimodal backbones. We propose RISCLIP, a novel framework that effectively leverages the cross-modal nature of CLIP for RIS. Observing CLIP's inherent alignment between image and text features, we capitalize on this starting point and introduce simple but strong modules that enhance unimodal feature extraction and leverage rich alignment knowledge in CLIP's image-text shared-embedding space. RISCLIP exhibits outstanding results on all three major RIS benchmarks and also outperforms previous CLIP-based methods, demonstrating the efficacy of our strategy in extending CLIP's image-text alignment to RIS.
title Extending CLIP's Image-Text Alignment to Referring Image Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2306.08498