LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866908314479099904 |
|---|---|
| author | Jin, Can Li, Ying Zhao, Mingyu Zhao, Shiyu Wang, Zhenting He, Xiaoxiao Han, Ligong Che, Tong Metaxas, Dimitris N. |
| author_facet | Jin, Can Li, Ying Zhao, Mingyu Zhao, Shiyu Wang, Zhenting He, Xiaoxiao Han, Ligong Che, Tong Metaxas, Dimitris N. |
| contents | Visual prompting has gained popularity as a method for adapting pre-trained models to specific tasks, particularly in the realm of parameter-efficient tuning. However, existing visual prompting techniques often pad the prompt parameters around the image, limiting the interaction between the visual prompts and the original image to a small set of patches while neglecting the inductive bias present in shared information across different patches. In this study, we conduct a thorough preliminary investigation to identify and address these limitations. We propose a novel visual prompt design, introducing Low-Rank matrix multiplication for Visual Prompting (LoR-VP), which enables shared and patch-specific information across rows and columns of image pixels. Extensive experiments across seven network architectures and four datasets demonstrate significant improvements in both performance and efficiency compared to state-of-the-art visual prompting methods, achieving up to 6 times faster training times, utilizing 18 times fewer visual prompt parameters, and delivering a 3.1% improvement in performance. The code is available as https://github.com/jincan333/LoR-VP. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_00896 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation Jin, Can Li, Ying Zhao, Mingyu Zhao, Shiyu Wang, Zhenting He, Xiaoxiao Han, Ligong Che, Tong Metaxas, Dimitris N. Computer Vision and Pattern Recognition Visual prompting has gained popularity as a method for adapting pre-trained models to specific tasks, particularly in the realm of parameter-efficient tuning. However, existing visual prompting techniques often pad the prompt parameters around the image, limiting the interaction between the visual prompts and the original image to a small set of patches while neglecting the inductive bias present in shared information across different patches. In this study, we conduct a thorough preliminary investigation to identify and address these limitations. We propose a novel visual prompt design, introducing Low-Rank matrix multiplication for Visual Prompting (LoR-VP), which enables shared and patch-specific information across rows and columns of image pixels. Extensive experiments across seven network architectures and four datasets demonstrate significant improvements in both performance and efficiency compared to state-of-the-art visual prompting methods, achieving up to 6 times faster training times, utilizing 18 times fewer visual prompt parameters, and delivering a 3.1% improvement in performance. The code is available as https://github.com/jincan333/LoR-VP. |
| title | LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2502.00896 |