Sparse-Tuning: Adapting Vision Transformers with Efficient Fine-tuning and Inference

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Ting, Liu, Xuyang, Shi, Liangtao, Xu, Zunnan, Hu, Yue, Huang, Siteng, Xin, Yi, Zhong, Bineng, Wang, Donglin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917152551862272
author Liu, Ting
Liu, Xuyang
Shi, Liangtao
Xu, Zunnan
Hu, Yue
Huang, Siteng
Xin, Yi
Zhong, Bineng
Wang, Donglin
author_facet Liu, Ting
Liu, Xuyang
Shi, Liangtao
Xu, Zunnan
Hu, Yue
Huang, Siteng
Xin, Yi
Zhong, Bineng
Wang, Donglin
contents Parameter-efficient fine-tuning (PEFT) has emerged as a popular solution for adapting pre-trained Vision Transformer (ViT) models to downstream applications by updating only a small subset of parameters. While current PEFT methods have achieved fine-tuning efficiency, they overlook the efficiency of computation and GPU memory during inference, falling short of practical requirements. To address this limitation, we propose Sparse-Tuning, an efficient and effective framework that leverages popular token sparsification (TS) techniques to reduce information redundancy in images and videos, thereby significantly improving computational and memory efficiency. However, TS often compromises performance due to inevitable information loss. To address this limitation, we further introduce Dense Adapters (DA) to compensate for the information losses incurred by token sparsification. DA integrates comprehensive token information from shallow layers into the retained tokens of deeper layers, ensuring minimal performance degradation. Through the integration of TS techniques and DA, Sparse-Tuning achieves a significant reduction in computation and memory overhead while maintaining performance. Empirical results on VTAB-1K, three image datasets, and two video datasets show that Sparse-Tuning reduces GFLOPs to 66\% of the original ViT-B while achieving state-of-the-art performance compared to full fine-tuning and other PEFT baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14700
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sparse-Tuning: Adapting Vision Transformers with Efficient Fine-tuning and Inference
Liu, Ting
Liu, Xuyang
Shi, Liangtao
Xu, Zunnan
Hu, Yue
Huang, Siteng
Xin, Yi
Zhong, Bineng
Wang, Donglin
Computer Vision and Pattern Recognition
Parameter-efficient fine-tuning (PEFT) has emerged as a popular solution for adapting pre-trained Vision Transformer (ViT) models to downstream applications by updating only a small subset of parameters. While current PEFT methods have achieved fine-tuning efficiency, they overlook the efficiency of computation and GPU memory during inference, falling short of practical requirements. To address this limitation, we propose Sparse-Tuning, an efficient and effective framework that leverages popular token sparsification (TS) techniques to reduce information redundancy in images and videos, thereby significantly improving computational and memory efficiency. However, TS often compromises performance due to inevitable information loss. To address this limitation, we further introduce Dense Adapters (DA) to compensate for the information losses incurred by token sparsification. DA integrates comprehensive token information from shallow layers into the retained tokens of deeper layers, ensuring minimal performance degradation. Through the integration of TS techniques and DA, Sparse-Tuning achieves a significant reduction in computation and memory overhead while maintaining performance. Empirical results on VTAB-1K, three image datasets, and two video datasets show that Sparse-Tuning reduces GFLOPs to 66\% of the original ViT-B while achieving state-of-the-art performance compared to full fine-tuning and other PEFT baselines.
title Sparse-Tuning: Adapting Vision Transformers with Efficient Fine-tuning and Inference
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.14700