AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zheng, Song, Yibing, Zhang, Xin, Luo, Lei, Li, Xiang, Yang, Jian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911288113758208
author Li, Zheng
Song, Yibing
Zhang, Xin
Luo, Lei
Li, Xiang
Yang, Jian
author_facet Li, Zheng
Song, Yibing
Zhang, Xin
Luo, Lei
Li, Xiang
Yang, Jian
contents Existing prompt learning methods, which are built upon CLIP models, leverage textual tokens as anchors to guide the learnable soft tokens. This guidance improves CLIP generalizations. However, these anchors-static in both value and position-lack cross-task and stage-adaptive flexibility. To address this limitation, we propose AnchorOPT, a dynamic anchor-based prompt learning framework. Specifically, AnchorOPT introduces dynamism in two key dimensions: (i) anchor values eschew handcrafted explicit textual tokens (e.g., "shape", "color"), instead learning dynamically from task-specific data; and (ii) the positional relationship between anchor and soft tokens is no longer fixed but adaptively optimized via a learnable position matrix conditioned on the training stage and task context. Training occurs in two stages: we first learn the anchor tokens, then freeze and transfer them to the second stage for optimization of soft tokens and the position matrix. Extensive experiments demonstrate that using only a simple learnable anchor and position matrix achieves performance comparable to or exceeding some methods incorporating additional learnable modules or regularization techniques. As a plug-and-play module, AnchorOPT integrates seamlessly into existing frameworks, yielding consistent performance gains across diverse datasets. Code is publicly available at https://github.com/zhengli97/ATPrompt.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21188
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning
Li, Zheng
Song, Yibing
Zhang, Xin
Luo, Lei
Li, Xiang
Yang, Jian
Computer Vision and Pattern Recognition
Computation and Language
Existing prompt learning methods, which are built upon CLIP models, leverage textual tokens as anchors to guide the learnable soft tokens. This guidance improves CLIP generalizations. However, these anchors-static in both value and position-lack cross-task and stage-adaptive flexibility. To address this limitation, we propose AnchorOPT, a dynamic anchor-based prompt learning framework. Specifically, AnchorOPT introduces dynamism in two key dimensions: (i) anchor values eschew handcrafted explicit textual tokens (e.g., "shape", "color"), instead learning dynamically from task-specific data; and (ii) the positional relationship between anchor and soft tokens is no longer fixed but adaptively optimized via a learnable position matrix conditioned on the training stage and task context. Training occurs in two stages: we first learn the anchor tokens, then freeze and transfer them to the second stage for optimization of soft tokens and the position matrix. Extensive experiments demonstrate that using only a simple learnable anchor and position matrix achieves performance comparable to or exceeding some methods incorporating additional learnable modules or regularization techniques. As a plug-and-play module, AnchorOPT integrates seamlessly into existing frameworks, yielding consistent performance gains across diverse datasets. Code is publicly available at https://github.com/zhengli97/ATPrompt.
title AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2511.21188