Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jinpeng, Luo, Tianci, Zha, Yaohua, Feng, Yan, Luo, Ruisheng, Chen, Bin, Dai, Tao, Chen, Long, Wang, Yaowei, Xia, Shu-Tao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
by: Luo, Tianci, et al.
Published: (2026)
by: Luo, Tianci, et al.
Published: (2026)
AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing
by: Lian, Niu, et al.
Published: (2025)
by: Lian, Niu, et al.
Published: (2025)
Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval
by: Li, Jun, et al.
Published: (2026)
by: Li, Jun, et al.
Published: (2026)
HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
by: Li, Jun, et al.
Published: (2025)
by: Li, Jun, et al.
Published: (2025)
GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video Retrieval
by: Wang, Yuting, et al.
Published: (2023)
by: Wang, Yuting, et al.
Published: (2023)
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
by: Lian, Niu, et al.
Published: (2026)
by: Lian, Niu, et al.
Published: (2026)
Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval
by: Li, Jun, et al.
Published: (2026)
by: Li, Jun, et al.
Published: (2026)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
by: Li, Liupeng, et al.
Published: (2026)
by: Li, Liupeng, et al.
Published: (2026)
CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
by: Lu, Zhenyu, et al.
Published: (2025)
by: Lu, Zhenyu, et al.
Published: (2025)
Efficient Self-Supervised Video Hashing with Selective State Spaces
by: Wang, Jinpeng, et al.
Published: (2024)
by: Wang, Jinpeng, et al.
Published: (2024)
Towards Efficient Low-rate Image Compression with Frequency-aware Diffusion Prior Refinement
by: Xia, Yichong, et al.
Published: (2026)
by: Xia, Yichong, et al.
Published: (2026)
PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and Alignment
by: Luo, Tianci, et al.
Published: (2026)
by: Luo, Tianci, et al.
Published: (2026)
AIM: Let Any Multi-modal Large Language Models Embrace Efficient In-Context Learning
by: Gao, Jun, et al.
Published: (2024)
by: Gao, Jun, et al.
Published: (2024)
GMMFormer v2: An Uncertainty-aware Framework for Partially Relevant Video Retrieval
by: Wang, Yuting, et al.
Published: (2024)
by: Wang, Yuting, et al.
Published: (2024)
3D-LMVIC: Learning-based Multi-View Image Coding with 3D Gaussian Geometric Priors
by: Huang, Yujun, et al.
Published: (2024)
by: Huang, Yujun, et al.
Published: (2024)
VP3D: Unleashing 2D Visual Prompt for Text-to-3D Generation
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Efficient Prompt Tuning for Hierarchical Ingredient Recognition
by: Gui, Yinxuan, et al.
Published: (2025)
by: Gui, Yinxuan, et al.
Published: (2025)
DPC: Dual-Prompt Collaboration for Tuning Vision-Language Models
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
ProMSC-MIS: Prompt-based Multimodal Semantic Communication for Multi-Spectral Image Segmentation
by: Zhang, Haoshuo, et al.
Published: (2025)
by: Zhang, Haoshuo, et al.
Published: (2025)
Wireless Multi-User Interactive Virtual Reality in Metaverse with Edge-Device Collaborative Computing
by: Xu, Caolu, et al.
Published: (2024)
by: Xu, Caolu, et al.
Published: (2024)
MambaVC: Learned Visual Compression with Selective State Spaces
by: Qin, Shiyu, et al.
Published: (2024)
by: Qin, Shiyu, et al.
Published: (2024)
ITEACH-Net: Inverted Teacher-studEnt seArCH Network for Emotion Recognition in Conversation
by: Sun, Haiyang, et al.
Published: (2023)
by: Sun, Haiyang, et al.
Published: (2023)
Interpreting Multimodal Communication at Scale in Short-Form Video: Visual, Audio, and Textual Mental Health Discourse on TikTok
by: Zha, Mingyue, et al.
Published: (2026)
by: Zha, Mingyue, et al.
Published: (2026)
Target Speech Diarization with Multimodal Prompts
by: Jiang, Yidi, et al.
Published: (2024)
by: Jiang, Yidi, et al.
Published: (2024)
AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
by: Chen, Zixuan, et al.
Published: (2026)
by: Chen, Zixuan, et al.
Published: (2026)
CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction
by: Wang, Jiadong, et al.
Published: (2026)
by: Wang, Jiadong, et al.
Published: (2026)
SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation
by: Lu, Zhenyu, et al.
Published: (2026)
by: Lu, Zhenyu, et al.
Published: (2026)
Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person Retrieval
by: Luo, Bingjun, et al.
Published: (2025)
by: Luo, Bingjun, et al.
Published: (2025)
Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution
by: Chen, Zhikai, et al.
Published: (2024)
by: Chen, Zhikai, et al.
Published: (2024)
Delayed Commitment for Representation Readiness in Stage-wise Audio-Visual Learning
by: Xu, Xinmeng, et al.
Published: (2026)
by: Xu, Xinmeng, et al.
Published: (2026)
EidetiCom: A Cross-modal Brain-Computer Semantic Communication Paradigm for Decoding Visual Perception
by: Zheng, Linfeng, et al.
Published: (2024)
by: Zheng, Linfeng, et al.
Published: (2024)
CLAIP-Emo: Parameter-Efficient Adaptation of Language-supervised models for In-the-Wild Audiovisual Emotion Recognition
by: Chen, Yin, et al.
Published: (2025)
by: Chen, Yin, et al.
Published: (2025)
Video Quality Assessment for Resolution Cross-Over in Live Sports
by: Zhu, Jingwen, et al.
Published: (2025)
by: Zhu, Jingwen, et al.
Published: (2025)
Stepwise Schema-Guided Prompting Framework with Parameter Efficient Instruction Tuning for Multimedia Event Extraction
by: Yuan, Xiang, et al.
Published: (2025)
by: Yuan, Xiang, et al.
Published: (2025)
MOMENTA: Mixture-of-Experts Over Multimodal Embeddings with Neural Temporal Aggregation for Misinformation Detection
by: Abdollahinejad, Yeganeh, et al.
Published: (2026)
by: Abdollahinejad, Yeganeh, et al.
Published: (2026)
MViR: Multi-View Visual-Semantic Representation for Fake News Detection
by: Liang, Haochen, et al.
Published: (2026)
by: Liang, Haochen, et al.
Published: (2026)
UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning
by: Bai, Hayes, et al.
Published: (2026)
by: Bai, Hayes, et al.
Published: (2026)
Multiple Contexts and Frequencies Aggregation Network forDeepfake Detection
by: Li, Zifeng, et al.
Published: (2024)
by: Li, Zifeng, et al.
Published: (2024)
Multimodal Emotion Recognition by Fusing Video Semantic in MOOC Learning Scenarios
by: Zhang, Yuan, et al.
Published: (2024)
by: Zhang, Yuan, et al.
Published: (2024)
VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
by: Chen, Siran, et al.
Published: (2025)
by: Chen, Siran, et al.
Published: (2025)
Similar Items
-
Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
by: Luo, Tianci, et al.
Published: (2026) -
AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing
by: Lian, Niu, et al.
Published: (2025) -
Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval
by: Li, Jun, et al.
Published: (2026) -
HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
by: Li, Jun, et al.
Published: (2025) -
GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video Retrieval
by: Wang, Yuting, et al.
Published: (2023)