PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Tianci, Wang, Jinpeng, Qin, Shiyu, Lian, Niu, Feng, Yan, Chen, Bin, Yuan, Chun, Xia, Shu-Tao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912974281637888
author Luo, Tianci
Wang, Jinpeng
Qin, Shiyu
Lian, Niu
Feng, Yan
Chen, Bin
Yuan, Chun
Xia, Shu-Tao
author_facet Luo, Tianci
Wang, Jinpeng
Qin, Shiyu
Lian, Niu
Feng, Yan
Chen, Bin
Yuan, Chun
Xia, Shu-Tao
contents Visual In-Context Learning (VICL) aims to complete vision tasks by imitating pixel demonstrations. Recent work pioneered prompt fusion that combines the advantages of various demonstrations, which shows a promising way to extend VICL. Unfortunately, the patch-wise fusion framework and model-agnostic supervision hinder the exploitation of informative cues, thereby limiting performance gains. To overcome this deficiency, we introduce PromptHub, a framework that holistically strengthens multi-prompting through locality-aware fusion, concentration and alignment. PromptHub exploits spatial priors to capture richer contextual information, employs complementary concentration, alignment, and prediction objectives to mutually guide training, and incorporates data augmentation to further reinforce supervision. Extensive experiments on three fundamental vision tasks demonstrate the superiority of PromptHub. Moreover, we validate its universality, transferability, and robustness across out-of-distribution settings, and various retrieval scenarios. This work establishes a reliable locality-aware paradigm for prompt fusion, moving beyond prior patch-wise approaches. Code is available at https://github.com/luotc-why/ICLR26-PromptHub.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18891
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and Alignment
Luo, Tianci
Wang, Jinpeng
Qin, Shiyu
Lian, Niu
Feng, Yan
Chen, Bin
Yuan, Chun
Xia, Shu-Tao
Computer Vision and Pattern Recognition
Machine Learning
Visual In-Context Learning (VICL) aims to complete vision tasks by imitating pixel demonstrations. Recent work pioneered prompt fusion that combines the advantages of various demonstrations, which shows a promising way to extend VICL. Unfortunately, the patch-wise fusion framework and model-agnostic supervision hinder the exploitation of informative cues, thereby limiting performance gains. To overcome this deficiency, we introduce PromptHub, a framework that holistically strengthens multi-prompting through locality-aware fusion, concentration and alignment. PromptHub exploits spatial priors to capture richer contextual information, employs complementary concentration, alignment, and prediction objectives to mutually guide training, and incorporates data augmentation to further reinforce supervision. Extensive experiments on three fundamental vision tasks demonstrate the superiority of PromptHub. Moreover, we validate its universality, transferability, and robustness across out-of-distribution settings, and various retrieval scenarios. This work establishes a reliable locality-aware paradigm for prompt fusion, moving beyond prior patch-wise approaches. Code is available at https://github.com/luotc-why/ICLR26-PromptHub.
title PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and Alignment
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2603.18891