TAI++: Text as Image for Multi-Label Image Classification by Co-Learning Transferable Prompt

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Xiangyu, Jiang, Qing-Yuan, Yang, Yang, Wu, Yi-Feng, Chen, Qing-Guo, Lu, Jianfeng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914793141567488
author Wu, Xiangyu
Jiang, Qing-Yuan
Yang, Yang
Wu, Yi-Feng
Chen, Qing-Guo
Lu, Jianfeng
author_facet Wu, Xiangyu
Jiang, Qing-Yuan
Yang, Yang
Wu, Yi-Feng
Chen, Qing-Guo
Lu, Jianfeng
contents The recent introduction of prompt tuning based on pre-trained vision-language models has dramatically improved the performance of multi-label image classification. However, some existing strategies that have been explored still have drawbacks, i.e., either exploiting massive labeled visual data at a high cost or using text data only for text prompt tuning and thus failing to learn the diversity of visual knowledge. Hence, the application scenarios of these methods are limited. In this paper, we propose a pseudo-visual prompt~(PVP) module for implicit visual prompt tuning to address this problem. Specifically, we first learn the pseudo-visual prompt for each category, mining diverse visual knowledge by the well-aligned space of pre-trained vision-language models. Then, a co-learning strategy with a dual-adapter module is designed to transfer visual knowledge from pseudo-visual prompt to text prompt, enhancing their visual representation abilities. Experimental results on VOC2007, MS-COCO, and NUSWIDE datasets demonstrate that our method can surpass state-of-the-art~(SOTA) methods across various settings for multi-label image classification tasks. The code is available at https://github.com/njustkmg/PVP.
format Preprint
id arxiv_https___arxiv_org_abs_2405_06926
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TAI++: Text as Image for Multi-Label Image Classification by Co-Learning Transferable Prompt
Wu, Xiangyu
Jiang, Qing-Yuan
Yang, Yang
Wu, Yi-Feng
Chen, Qing-Guo
Lu, Jianfeng
Computer Vision and Pattern Recognition
The recent introduction of prompt tuning based on pre-trained vision-language models has dramatically improved the performance of multi-label image classification. However, some existing strategies that have been explored still have drawbacks, i.e., either exploiting massive labeled visual data at a high cost or using text data only for text prompt tuning and thus failing to learn the diversity of visual knowledge. Hence, the application scenarios of these methods are limited. In this paper, we propose a pseudo-visual prompt~(PVP) module for implicit visual prompt tuning to address this problem. Specifically, we first learn the pseudo-visual prompt for each category, mining diverse visual knowledge by the well-aligned space of pre-trained vision-language models. Then, a co-learning strategy with a dual-adapter module is designed to transfer visual knowledge from pseudo-visual prompt to text prompt, enhancing their visual representation abilities. Experimental results on VOC2007, MS-COCO, and NUSWIDE datasets demonstrate that our method can surpass state-of-the-art~(SOTA) methods across various settings for multi-label image classification tasks. The code is available at https://github.com/njustkmg/PVP.
title TAI++: Text as Image for Multi-Label Image Classification by Co-Learning Transferable Prompt
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.06926