Towards Efficient Vision-Language Tuning: More Information Density, More Generalizability

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hao, Tianxiang, Lyu, Mengyao, Chen, Hui, Zhao, Sicheng, Ding, Xiaohan, Han, Jungong, Ding, Guiguang
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917772519276544
author Hao, Tianxiang
Lyu, Mengyao
Chen, Hui
Zhao, Sicheng
Ding, Xiaohan
Han, Jungong
Ding, Guiguang
author_facet Hao, Tianxiang
Lyu, Mengyao
Chen, Hui
Zhao, Sicheng
Ding, Xiaohan
Han, Jungong
Ding, Guiguang
contents With the advancement of large pre-trained vision-language models, effectively transferring the knowledge embedded within these foundational models to downstream tasks has become a pivotal topic, particularly in data-scarce environments. Recently, parameter-efficient fine-tuning approaches, especially prompt tuning, have garnered considerable attention. To better understand the nature of prompt tuning, we propose the concept of ``Information Density'' (ID) to indicate whether a matrix strongly belongs to certain feature spaces rather than being evenly distributed across various feature spaces. We suppose a higher ID with strong bias across some feature spaces naturally leads to excellent robustness and stability. Our research, inspired by the observation that generalizability is closely linked to the information density of the prompt matrix, introduces the Dense Information Prompt (DIP). DIP aims to enhance information density to improve generalization. Furthermore, DIP significantly reduces the number of tunable parameters and the requisite storage space, making it particularly advantageous in resource-constrained settings. Comprehensive experiments substantiate the superiority of DIP. Notably, DIP surpasses the latest state-of-the-art methods by a substantial margin with an exceptionally small parameter count. Across a range of tasks spanning 11 datasets, DIP improves the average downstream accuracy of classic prompt tuning by up to 5.76% using merely 0.5K parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2312_10813
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Towards Efficient Vision-Language Tuning: More Information Density, More Generalizability
Hao, Tianxiang
Lyu, Mengyao
Chen, Hui
Zhao, Sicheng
Ding, Xiaohan
Han, Jungong
Ding, Guiguang
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
With the advancement of large pre-trained vision-language models, effectively transferring the knowledge embedded within these foundational models to downstream tasks has become a pivotal topic, particularly in data-scarce environments. Recently, parameter-efficient fine-tuning approaches, especially prompt tuning, have garnered considerable attention. To better understand the nature of prompt tuning, we propose the concept of ``Information Density'' (ID) to indicate whether a matrix strongly belongs to certain feature spaces rather than being evenly distributed across various feature spaces. We suppose a higher ID with strong bias across some feature spaces naturally leads to excellent robustness and stability. Our research, inspired by the observation that generalizability is closely linked to the information density of the prompt matrix, introduces the Dense Information Prompt (DIP). DIP aims to enhance information density to improve generalization. Furthermore, DIP significantly reduces the number of tunable parameters and the requisite storage space, making it particularly advantageous in resource-constrained settings. Comprehensive experiments substantiate the superiority of DIP. Notably, DIP surpasses the latest state-of-the-art methods by a substantial margin with an exceptionally small parameter count. Across a range of tasks spanning 11 datasets, DIP improves the average downstream accuracy of classic prompt tuning by up to 5.76% using merely 0.5K parameters.
title Towards Efficient Vision-Language Tuning: More Information Density, More Generalizability
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2312.10813