Efficient and Effective In-context Demonstration Selection with Coreset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zihua, Wang, Jiarui, Xu, Haiyang, Yan, Ming, Huang, Fei, Yang, Xu, Wei, Xiu-Shen, Mi, Siya, Zhang, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917075768836096
author Wang, Zihua
Wang, Jiarui
Xu, Haiyang
Yan, Ming
Huang, Fei
Yang, Xu
Wei, Xiu-Shen
Mi, Siya
Zhang, Yu
author_facet Wang, Zihua
Wang, Jiarui
Xu, Haiyang
Yan, Ming
Huang, Fei
Yang, Xu
Wei, Xiu-Shen
Mi, Siya
Zhang, Yu
contents In-context learning (ICL) has emerged as a powerful paradigm for Large Visual Language Models (LVLMs), enabling them to leverage a few examples directly from input contexts. However, the effectiveness of this approach is heavily reliant on the selection of demonstrations, a process that is NP-hard. Traditional strategies, including random, similarity-based sampling and infoscore-based sampling, often lead to inefficiencies or suboptimal performance, struggling to balance both efficiency and effectiveness in demonstration selection. In this paper, we propose a novel demonstration selection framework named Coreset-based Dual Retrieval (CoDR). We show that samples within a diverse subset achieve a higher expected mutual information. To implement this, we introduce a cluster-pruning method to construct a diverse coreset that aligns more effectively with the query while maintaining diversity. Additionally, we develop a dual retrieval mechanism that enhances the selection process by achieving global demonstration selection while preserving efficiency. Experimental results demonstrate that our method significantly improves the ICL performance compared to the existing strategies, providing a robust solution for effective and efficient demonstration selection.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08977
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient and Effective In-context Demonstration Selection with Coreset
Wang, Zihua
Wang, Jiarui
Xu, Haiyang
Yan, Ming
Huang, Fei
Yang, Xu
Wei, Xiu-Shen
Mi, Siya
Zhang, Yu
Computer Vision and Pattern Recognition
In-context learning (ICL) has emerged as a powerful paradigm for Large Visual Language Models (LVLMs), enabling them to leverage a few examples directly from input contexts. However, the effectiveness of this approach is heavily reliant on the selection of demonstrations, a process that is NP-hard. Traditional strategies, including random, similarity-based sampling and infoscore-based sampling, often lead to inefficiencies or suboptimal performance, struggling to balance both efficiency and effectiveness in demonstration selection. In this paper, we propose a novel demonstration selection framework named Coreset-based Dual Retrieval (CoDR). We show that samples within a diverse subset achieve a higher expected mutual information. To implement this, we introduce a cluster-pruning method to construct a diverse coreset that aligns more effectively with the query while maintaining diversity. Additionally, we develop a dual retrieval mechanism that enhances the selection process by achieving global demonstration selection while preserving efficiency. Experimental results demonstrate that our method significantly improves the ICL performance compared to the existing strategies, providing a robust solution for effective and efficient demonstration selection.
title Efficient and Effective In-context Demonstration Selection with Coreset
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.08977