Context-Based Semantic-Aware Alignment for Semi-Supervised Multi-Label Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fan, Heng-Bo, Xie, Ming-Kun, Xiao, Jia-Hao, Huang, Sheng-Jun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913627013906432
author Fan, Heng-Bo
Xie, Ming-Kun
Xiao, Jia-Hao
Huang, Sheng-Jun
author_facet Fan, Heng-Bo
Xie, Ming-Kun
Xiao, Jia-Hao
Huang, Sheng-Jun
contents Due to the lack of extensive precisely-annotated multi-label data in real word, semi-supervised multi-label learning (SSMLL) has gradually gained attention. Abundant knowledge embedded in vision-language models (VLMs) pre-trained on large-scale image-text pairs could alleviate the challenge of limited labeled data under SSMLL setting.Despite existing methods based on fine-tuning VLMs have achieved advances in weakly-supervised multi-label learning, they failed to fully leverage the information from labeled data to enhance the learning of unlabeled data. In this paper, we propose a context-based semantic-aware alignment method to solve the SSMLL problem by leveraging the knowledge of VLMs. To address the challenge of handling multiple semantics within an image, we introduce a novel framework design to extract label-specific image features. This design allows us to achieve a more compact alignment between text features and label-specific image features, leading the model to generate high-quality pseudo-labels. To incorporate the model with comprehensive understanding of image, we design a semi-supervised context identification auxiliary task to enhance the feature representation by capturing co-occurrence information. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness of our proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18842
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Context-Based Semantic-Aware Alignment for Semi-Supervised Multi-Label Learning
Fan, Heng-Bo
Xie, Ming-Kun
Xiao, Jia-Hao
Huang, Sheng-Jun
Computer Vision and Pattern Recognition
Machine Learning
Due to the lack of extensive precisely-annotated multi-label data in real word, semi-supervised multi-label learning (SSMLL) has gradually gained attention. Abundant knowledge embedded in vision-language models (VLMs) pre-trained on large-scale image-text pairs could alleviate the challenge of limited labeled data under SSMLL setting.Despite existing methods based on fine-tuning VLMs have achieved advances in weakly-supervised multi-label learning, they failed to fully leverage the information from labeled data to enhance the learning of unlabeled data. In this paper, we propose a context-based semantic-aware alignment method to solve the SSMLL problem by leveraging the knowledge of VLMs. To address the challenge of handling multiple semantics within an image, we introduce a novel framework design to extract label-specific image features. This design allows us to achieve a more compact alignment between text features and label-specific image features, leading the model to generate high-quality pseudo-labels. To incorporate the model with comprehensive understanding of image, we design a semi-supervised context identification auxiliary task to enhance the feature representation by capturing co-occurrence information. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness of our proposed method.
title Context-Based Semantic-Aware Alignment for Semi-Supervised Multi-Label Learning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2412.18842