Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIP

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Zhongxing, Tang, Feilong, Chen, Zhe, Su, Yingxue, Zhao, Zhiyi, Zhang, Ge, Su, Jionglong, Ge, Zongyuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912170490462208
author Xu, Zhongxing
Tang, Feilong
Chen, Zhe
Su, Yingxue
Zhao, Zhiyi
Zhang, Ge
Su, Jionglong
Ge, Zongyuan
author_facet Xu, Zhongxing
Tang, Feilong
Chen, Zhe
Su, Yingxue
Zhao, Zhiyi
Zhang, Ge
Su, Jionglong
Ge, Zongyuan
contents The application of Contrastive Language-Image Pre-training (CLIP) in Weakly Supervised Semantic Segmentation (WSSS) research powerful cross-modal semantic understanding capabilities. Existing methods attempt to optimize input text prompts for improved alignment of images and text, by finely adjusting text prototypes to facilitate semantic matching. Nevertheless, given the modality gap between text and vision spaces, the text prototypes employed by these methods have not effectively established a close correspondence with pixel-level vision features. In this work, our theoretical analysis indicates that the inherent modality gap results in misalignment of text and region features, and that this gap cannot be sufficiently reduced by minimizing contrast loss in CLIP. To mitigate the impact of the modality gap, we propose a Vision Prototype Learning (VPL) framework, by introducing more representative vision prototypes. The core of this framework is to learn class-specific vision prototypes in vision space with the help of text prototypes, for capturing high-quality localization maps. Moreover, we propose a regional semantic contrast module that contrasts regions embedding with corresponding prototypes, leading to more comprehensive and robust feature learning. Experimental results show that our proposed framework achieves state-of-the-art performance on two benchmark datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19650
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIP
Xu, Zhongxing
Tang, Feilong
Chen, Zhe
Su, Yingxue
Zhao, Zhiyi
Zhang, Ge
Su, Jionglong
Ge, Zongyuan
Computer Vision and Pattern Recognition
Machine Learning
The application of Contrastive Language-Image Pre-training (CLIP) in Weakly Supervised Semantic Segmentation (WSSS) research powerful cross-modal semantic understanding capabilities. Existing methods attempt to optimize input text prompts for improved alignment of images and text, by finely adjusting text prototypes to facilitate semantic matching. Nevertheless, given the modality gap between text and vision spaces, the text prototypes employed by these methods have not effectively established a close correspondence with pixel-level vision features. In this work, our theoretical analysis indicates that the inherent modality gap results in misalignment of text and region features, and that this gap cannot be sufficiently reduced by minimizing contrast loss in CLIP. To mitigate the impact of the modality gap, we propose a Vision Prototype Learning (VPL) framework, by introducing more representative vision prototypes. The core of this framework is to learn class-specific vision prototypes in vision space with the help of text prototypes, for capturing high-quality localization maps. Moreover, we propose a regional semantic contrast module that contrasts regions embedding with corresponding prototypes, leading to more comprehensive and robust feature learning. Experimental results show that our proposed framework achieves state-of-the-art performance on two benchmark datasets.
title Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIP
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2412.19650