Prompt-based Adaptation in Large-scale Vision Models: A Survey

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xiao, Xi, Zhang, Yunbei, Zhao, Lin, Liu, Yiyang, Liao, Xiaoying, Mai, Zheda, Li, Xingjian, Wang, Xiao, Xu, Hao, Hamm, Jihun, Lin, Xue, Xu, Min, Wang, Qifan, Wang, Tianyang, Han, Cheng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912974734622720
author Xiao, Xi
Zhang, Yunbei
Zhao, Lin
Liu, Yiyang
Liao, Xiaoying
Mai, Zheda
Li, Xingjian
Wang, Xiao
Xu, Hao
Hamm, Jihun
Lin, Xue
Xu, Min
Wang, Qifan
Wang, Tianyang
Han, Cheng
author_facet Xiao, Xi
Zhang, Yunbei
Zhao, Lin
Liu, Yiyang
Liao, Xiaoying
Mai, Zheda
Li, Xingjian
Wang, Xiao
Xu, Hao
Hamm, Jihun
Lin, Xue
Xu, Min
Wang, Qifan
Wang, Tianyang
Han, Cheng
contents In computer vision, Visual Prompting (VP) and Visual Prompt Tuning (VPT) have recently emerged as lightweight and effective alternatives to full fine-tuning for adapting large-scale vision models within the "pretrain-then-finetune" paradigm. However, despite rapid progress, their conceptual boundaries remain blurred, as VP and VPT are frequently used interchangeably in current research, reflecting a lack of systematic distinction between these techniques and their respective applications. In this survey, we revisit the designs of VP and VPT from first principles and conceptualize them within a unified framework termed Prompt-based Adaptation (PA). Within this framework, we distinguish methods based on their injection granularity: VP operates at the pixel level, while VPT injects prompts at the token level. We further categorize these methods by their generation mechanism into fixed, learnable, and generated prompts. Beyond the core methodologies, we examine PA integrations across diverse domains, including medical imaging, 3D point clouds, and vision-language tasks, as well as its role in test-time adaptation and trustworthy AI. We also summarize current benchmarks and identify key challenges and future directions. To the best of our knowledge, we are the first comprehensive survey dedicated to PA methodologies and applications in light of their distinct characteristics. Our survey aims to provide a clear roadmap for researchers and practitioners in all areas to understand and explore the evolving landscape of PA-related research.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13219
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prompt-based Adaptation in Large-scale Vision Models: A Survey
Xiao, Xi
Zhang, Yunbei
Zhao, Lin
Liu, Yiyang
Liao, Xiaoying
Mai, Zheda
Li, Xingjian
Wang, Xiao
Xu, Hao
Hamm, Jihun
Lin, Xue
Xu, Min
Wang, Qifan
Wang, Tianyang
Han, Cheng
Computer Vision and Pattern Recognition
In computer vision, Visual Prompting (VP) and Visual Prompt Tuning (VPT) have recently emerged as lightweight and effective alternatives to full fine-tuning for adapting large-scale vision models within the "pretrain-then-finetune" paradigm. However, despite rapid progress, their conceptual boundaries remain blurred, as VP and VPT are frequently used interchangeably in current research, reflecting a lack of systematic distinction between these techniques and their respective applications. In this survey, we revisit the designs of VP and VPT from first principles and conceptualize them within a unified framework termed Prompt-based Adaptation (PA). Within this framework, we distinguish methods based on their injection granularity: VP operates at the pixel level, while VPT injects prompts at the token level. We further categorize these methods by their generation mechanism into fixed, learnable, and generated prompts. Beyond the core methodologies, we examine PA integrations across diverse domains, including medical imaging, 3D point clouds, and vision-language tasks, as well as its role in test-time adaptation and trustworthy AI. We also summarize current benchmarks and identify key challenges and future directions. To the best of our knowledge, we are the first comprehensive survey dedicated to PA methodologies and applications in light of their distinct characteristics. Our survey aims to provide a clear roadmap for researchers and practitioners in all areas to understand and explore the evolving landscape of PA-related research.
title Prompt-based Adaptation in Large-scale Vision Models: A Survey
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.13219