LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jin, Can, Li, Ying, Zhao, Mingyu, Zhao, Shiyu, Wang, Zhenting, He, Xiaoxiao, Han, Ligong, Che, Tong, Metaxas, Dimitris N.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908314479099904
author Jin, Can
Li, Ying
Zhao, Mingyu
Zhao, Shiyu
Wang, Zhenting
He, Xiaoxiao
Han, Ligong
Che, Tong
Metaxas, Dimitris N.
author_facet Jin, Can
Li, Ying
Zhao, Mingyu
Zhao, Shiyu
Wang, Zhenting
He, Xiaoxiao
Han, Ligong
Che, Tong
Metaxas, Dimitris N.
contents Visual prompting has gained popularity as a method for adapting pre-trained models to specific tasks, particularly in the realm of parameter-efficient tuning. However, existing visual prompting techniques often pad the prompt parameters around the image, limiting the interaction between the visual prompts and the original image to a small set of patches while neglecting the inductive bias present in shared information across different patches. In this study, we conduct a thorough preliminary investigation to identify and address these limitations. We propose a novel visual prompt design, introducing Low-Rank matrix multiplication for Visual Prompting (LoR-VP), which enables shared and patch-specific information across rows and columns of image pixels. Extensive experiments across seven network architectures and four datasets demonstrate significant improvements in both performance and efficiency compared to state-of-the-art visual prompting methods, achieving up to 6 times faster training times, utilizing 18 times fewer visual prompt parameters, and delivering a 3.1% improvement in performance. The code is available as https://github.com/jincan333/LoR-VP.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00896
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
Jin, Can
Li, Ying
Zhao, Mingyu
Zhao, Shiyu
Wang, Zhenting
He, Xiaoxiao
Han, Ligong
Che, Tong
Metaxas, Dimitris N.
Computer Vision and Pattern Recognition
Visual prompting has gained popularity as a method for adapting pre-trained models to specific tasks, particularly in the realm of parameter-efficient tuning. However, existing visual prompting techniques often pad the prompt parameters around the image, limiting the interaction between the visual prompts and the original image to a small set of patches while neglecting the inductive bias present in shared information across different patches. In this study, we conduct a thorough preliminary investigation to identify and address these limitations. We propose a novel visual prompt design, introducing Low-Rank matrix multiplication for Visual Prompting (LoR-VP), which enables shared and patch-specific information across rows and columns of image pixels. Extensive experiments across seven network architectures and four datasets demonstrate significant improvements in both performance and efficiency compared to state-of-the-art visual prompting methods, achieving up to 6 times faster training times, utilizing 18 times fewer visual prompt parameters, and delivering a 3.1% improvement in performance. The code is available as https://github.com/jincan333/LoR-VP.
title LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.00896