Can Better Text Semantics in Prompt Tuning Improve VLM Generalization?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kuchibhotla, Hari Chandana, Kancheti, Sai Srinivas, Reddy, Abbavaram Gowtham, Balasubramanian, Vineeth N
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913397914730496
author Kuchibhotla, Hari Chandana
Kancheti, Sai Srinivas
Reddy, Abbavaram Gowtham
Balasubramanian, Vineeth N
author_facet Kuchibhotla, Hari Chandana
Kancheti, Sai Srinivas
Reddy, Abbavaram Gowtham
Balasubramanian, Vineeth N
contents Going beyond mere fine-tuning of vision-language models (VLMs), learnable prompt tuning has emerged as a promising, resource-efficient alternative. Despite their potential, effectively learning prompts faces the following challenges: (i) training in a low-shot scenario results in overfitting, limiting adaptability, and yielding weaker performance on newer classes or datasets; (ii) prompt-tuning's efficacy heavily relies on the label space, with decreased performance in large class spaces, signaling potential gaps in bridging image and class concepts. In this work, we investigate whether better text semantics can help address these concerns. In particular, we introduce a prompt-tuning method that leverages class descriptions obtained from Large Language Models (LLMs). These class descriptions are used to bridge image and text modalities. Our approach constructs part-level description-guided image and text features, which are subsequently aligned to learn more generalizable prompts. Our comprehensive experiments conducted across 11 benchmark datasets show that our method outperforms established methods, demonstrating substantial improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2405_07921
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can Better Text Semantics in Prompt Tuning Improve VLM Generalization?
Kuchibhotla, Hari Chandana
Kancheti, Sai Srinivas
Reddy, Abbavaram Gowtham
Balasubramanian, Vineeth N
Computer Vision and Pattern Recognition
Going beyond mere fine-tuning of vision-language models (VLMs), learnable prompt tuning has emerged as a promising, resource-efficient alternative. Despite their potential, effectively learning prompts faces the following challenges: (i) training in a low-shot scenario results in overfitting, limiting adaptability, and yielding weaker performance on newer classes or datasets; (ii) prompt-tuning's efficacy heavily relies on the label space, with decreased performance in large class spaces, signaling potential gaps in bridging image and class concepts. In this work, we investigate whether better text semantics can help address these concerns. In particular, we introduce a prompt-tuning method that leverages class descriptions obtained from Large Language Models (LLMs). These class descriptions are used to bridge image and text modalities. Our approach constructs part-level description-guided image and text features, which are subsequently aligned to learn more generalizable prompts. Our comprehensive experiments conducted across 11 benchmark datasets show that our method outperforms established methods, demonstrating substantial improvements.
title Can Better Text Semantics in Prompt Tuning Improve VLM Generalization?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.07921