MAO: Efficient Model-Agnostic Optimization of Prompt Tuning for Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Haoyang, Zhou, Siyu, Wang, Liang, Long, Guodong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915211365056512
author Li, Haoyang
Zhou, Siyu
Wang, Liang
Long, Guodong
author_facet Li, Haoyang
Zhou, Siyu
Wang, Liang
Long, Guodong
contents Though CLIP-based prompt tuning significantly enhances pre-trained Vision-Language Models, existing research focuses on reconstructing the model architecture, e.g., additional loss calculation and meta-networks. These approaches generally lead to increased complexity and extended training cost. To maintain the efficiency of the tuning process, we propose plug-and-play Model-Agnostic Optimization (MAO) for prompt tuning. Without altering any components of the prompt tuning backbone, we introduce a Data-Driven Enhancement framework to optimize the distribution of the initial data, and incorporate an Alterable Regularization module to boost the task-specific feature processing pipeline, thereby improving overall performance while maintaining low computational cost. Extensive experiments on MAO demonstrate its outstanding performance and efficiency. The code of MAO is available at: https://github.com/JREion/M.A.O .
format Preprint
id arxiv_https___arxiv_org_abs_2503_18160
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MAO: Efficient Model-Agnostic Optimization of Prompt Tuning for Vision-Language Models
Li, Haoyang
Zhou, Siyu
Wang, Liang
Long, Guodong
Computer Vision and Pattern Recognition
Multimedia
Though CLIP-based prompt tuning significantly enhances pre-trained Vision-Language Models, existing research focuses on reconstructing the model architecture, e.g., additional loss calculation and meta-networks. These approaches generally lead to increased complexity and extended training cost. To maintain the efficiency of the tuning process, we propose plug-and-play Model-Agnostic Optimization (MAO) for prompt tuning. Without altering any components of the prompt tuning backbone, we introduce a Data-Driven Enhancement framework to optimize the distribution of the initial data, and incorporate an Alterable Regularization module to boost the task-specific feature processing pipeline, thereby improving overall performance while maintaining low computational cost. Extensive experiments on MAO demonstrate its outstanding performance and efficiency. The code of MAO is available at: https://github.com/JREion/M.A.O .
title MAO: Efficient Model-Agnostic Optimization of Prompt Tuning for Vision-Language Models
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2503.18160