Learning Semantic-Aware Representation in Visual-Language Models for Multi-Label Recognition with Partial Labels

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ruan, Haoxian, Xu, Zhihua, Yang, Zhijing, Lu, Yongyi, Qin, Jinghui, Chen, Tianshui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912156494069760
author Ruan, Haoxian
Xu, Zhihua
Yang, Zhijing
Lu, Yongyi
Qin, Jinghui
Chen, Tianshui
author_facet Ruan, Haoxian
Xu, Zhihua
Yang, Zhijing
Lu, Yongyi
Qin, Jinghui
Chen, Tianshui
contents Multi-label recognition with partial labels (MLR-PL), in which only some labels are known while others are unknown for each image, is a practical task in computer vision, since collecting large-scale and complete multi-label datasets is difficult in real application scenarios. Recently, vision language models (e.g. CLIP) have demonstrated impressive transferability to downstream tasks in data limited or label limited settings. However, current CLIP-based methods suffer from semantic confusion in MLR task due to the lack of fine-grained information in the single global visual and textual representation for all categories. In this work, we address this problem by introducing a semantic decoupling module and a category-specific prompt optimization method in CLIP-based framework. Specifically, the semantic decoupling module following the visual encoder learns category-specific feature maps by utilizing the semantic-guided spatial attention mechanism. Moreover, the category-specific prompt optimization method is introduced to learn text representations aligned with category semantics. Therefore, the prediction of each category is independent, which alleviate the semantic confusion problem. Extensive experiments on Microsoft COCO 2014 and Pascal VOC 2007 datasets demonstrate that the proposed framework significantly outperforms current state-of-art methods with a simpler model structure. Additionally, visual analysis shows that our method effectively separates information from different categories and achieves better performance compared to CLIP-based baseline method.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10843
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Semantic-Aware Representation in Visual-Language Models for Multi-Label Recognition with Partial Labels
Ruan, Haoxian
Xu, Zhihua
Yang, Zhijing
Lu, Yongyi
Qin, Jinghui
Chen, Tianshui
Computer Vision and Pattern Recognition
Multi-label recognition with partial labels (MLR-PL), in which only some labels are known while others are unknown for each image, is a practical task in computer vision, since collecting large-scale and complete multi-label datasets is difficult in real application scenarios. Recently, vision language models (e.g. CLIP) have demonstrated impressive transferability to downstream tasks in data limited or label limited settings. However, current CLIP-based methods suffer from semantic confusion in MLR task due to the lack of fine-grained information in the single global visual and textual representation for all categories. In this work, we address this problem by introducing a semantic decoupling module and a category-specific prompt optimization method in CLIP-based framework. Specifically, the semantic decoupling module following the visual encoder learns category-specific feature maps by utilizing the semantic-guided spatial attention mechanism. Moreover, the category-specific prompt optimization method is introduced to learn text representations aligned with category semantics. Therefore, the prediction of each category is independent, which alleviate the semantic confusion problem. Extensive experiments on Microsoft COCO 2014 and Pascal VOC 2007 datasets demonstrate that the proposed framework significantly outperforms current state-of-art methods with a simpler model structure. Additionally, visual analysis shows that our method effectively separates information from different categories and achieves better performance compared to CLIP-based baseline method.
title Learning Semantic-Aware Representation in Visual-Language Models for Multi-Label Recognition with Partial Labels
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.10843