Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Zhiwei, Meng, Yucong, Fu, Kexue, Tang, Feilong, Wang, Shuo, Song, Zhijian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908319852003328
author Yang, Zhiwei
Meng, Yucong
Fu, Kexue
Tang, Feilong
Wang, Shuo
Song, Zhijian
author_facet Yang, Zhiwei
Meng, Yucong
Fu, Kexue
Tang, Feilong
Wang, Shuo
Song, Zhijian
contents Weakly Supervised Semantic Segmentation (WSSS) with image-level labels aims to achieve pixel-level predictions using Class Activation Maps (CAMs). Recently, Contrastive Language-Image Pre-training (CLIP) has been introduced in WSSS. However, recent methods primarily focus on image-text alignment for CAM generation, while CLIP's potential in patch-text alignment remains unexplored. In this work, we propose ExCEL to explore CLIP's dense knowledge via a novel patch-text alignment paradigm for WSSS. Specifically, we propose Text Semantic Enrichment (TSE) and Visual Calibration (VC) modules to improve the dense alignment across both text and vision modalities. To make text embeddings semantically informative, our TSE module applies Large Language Models (LLMs) to build a dataset-wide knowledge base and enriches the text representations with an implicit attribute-hunting process. To mine fine-grained knowledge from visual features, our VC module first proposes Static Visual Calibration (SVC) to propagate fine-grained knowledge in a non-parametric manner. Then Learnable Visual Calibration (LVC) is further proposed to dynamically shift the frozen features towards distributions with diverse semantics. With these enhancements, ExCEL not only retains CLIP's training-free advantages but also significantly outperforms other state-of-the-art methods with much less training cost on PASCAL VOC and MS COCO.
format Preprint
id arxiv_https___arxiv_org_abs_2503_20826
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation
Yang, Zhiwei
Meng, Yucong
Fu, Kexue
Tang, Feilong
Wang, Shuo
Song, Zhijian
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Image and Video Processing
Weakly Supervised Semantic Segmentation (WSSS) with image-level labels aims to achieve pixel-level predictions using Class Activation Maps (CAMs). Recently, Contrastive Language-Image Pre-training (CLIP) has been introduced in WSSS. However, recent methods primarily focus on image-text alignment for CAM generation, while CLIP's potential in patch-text alignment remains unexplored. In this work, we propose ExCEL to explore CLIP's dense knowledge via a novel patch-text alignment paradigm for WSSS. Specifically, we propose Text Semantic Enrichment (TSE) and Visual Calibration (VC) modules to improve the dense alignment across both text and vision modalities. To make text embeddings semantically informative, our TSE module applies Large Language Models (LLMs) to build a dataset-wide knowledge base and enriches the text representations with an implicit attribute-hunting process. To mine fine-grained knowledge from visual features, our VC module first proposes Static Visual Calibration (SVC) to propagate fine-grained knowledge in a non-parametric manner. Then Learnable Visual Calibration (LVC) is further proposed to dynamically shift the frozen features towards distributions with diverse semantics. With these enhancements, ExCEL not only retains CLIP's training-free advantages but also significantly outperforms other state-of-the-art methods with much less training cost on PASCAL VOC and MS COCO.
title Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Image and Video Processing
url https://arxiv.org/abs/2503.20826