Semantic Hierarchical Prompt Tuning for Parameter-Efficient Fine-Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Haowei, Zhang, Fangyuan, Qin, Rui, Pan, Tianxiang, Yong, Junhai, Wang, Bin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916541196402688
author Zhu, Haowei
Zhang, Fangyuan
Qin, Rui
Pan, Tianxiang
Yong, Junhai
Wang, Bin
author_facet Zhu, Haowei
Zhang, Fangyuan
Qin, Rui
Pan, Tianxiang
Yong, Junhai
Wang, Bin
contents As the scale of vision models continues to grow, Visual Prompt Tuning (VPT) has emerged as a parameter-efficient transfer learning technique, noted for its superior performance compared to full fine-tuning. However, indiscriminately applying prompts to every layer without considering their inherent correlations, can cause significant disturbances, leading to suboptimal transferability. Additionally, VPT disrupts the original self-attention structure, affecting the aggregation of visual features, and lacks a mechanism for explicitly mining discriminative visual features, which are crucial for classification. To address these issues, we propose a Semantic Hierarchical Prompt (SHIP) fine-tuning strategy. We adaptively construct semantic hierarchies and use semantic-independent and semantic-shared prompts to learn hierarchical representations. We also integrate attribute prompts and a prompt matching loss to enhance feature discrimination and employ decoupled attention for robustness and reduced inference costs. SHIP significantly improves performance, achieving a 4.9% gain in accuracy over VPT with a ViT-B/16 backbone on VTAB-1k tasks. Our code is available at https://github.com/haoweiz23/SHIP.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16956
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Semantic Hierarchical Prompt Tuning for Parameter-Efficient Fine-Tuning
Zhu, Haowei
Zhang, Fangyuan
Qin, Rui
Pan, Tianxiang
Yong, Junhai
Wang, Bin
Computer Vision and Pattern Recognition
As the scale of vision models continues to grow, Visual Prompt Tuning (VPT) has emerged as a parameter-efficient transfer learning technique, noted for its superior performance compared to full fine-tuning. However, indiscriminately applying prompts to every layer without considering their inherent correlations, can cause significant disturbances, leading to suboptimal transferability. Additionally, VPT disrupts the original self-attention structure, affecting the aggregation of visual features, and lacks a mechanism for explicitly mining discriminative visual features, which are crucial for classification. To address these issues, we propose a Semantic Hierarchical Prompt (SHIP) fine-tuning strategy. We adaptively construct semantic hierarchies and use semantic-independent and semantic-shared prompts to learn hierarchical representations. We also integrate attribute prompts and a prompt matching loss to enhance feature discrimination and employ decoupled attention for robustness and reduced inference costs. SHIP significantly improves performance, achieving a 4.9% gain in accuracy over VPT with a ViT-B/16 backbone on VTAB-1k tasks. Our code is available at https://github.com/haoweiz23/SHIP.
title Semantic Hierarchical Prompt Tuning for Parameter-Efficient Fine-Tuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.16956