Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Cheng, Zhai, Yunpeng, Zhao, Yifan, Gao, Jinyang, Ding, Bolin, Li, Jia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911000090902528
author Chen, Cheng
Zhai, Yunpeng
Zhao, Yifan
Gao, Jinyang
Ding, Bolin
Li, Jia
author_facet Chen, Cheng
Zhai, Yunpeng
Zhao, Yifan
Gao, Jinyang
Ding, Bolin
Li, Jia
contents In-context learning (ICL), a predominant trend in instruction learning, aims at enhancing the performance of large language models by providing clear task guidance and examples, improving their capability in task understanding and execution. This paper investigates ICL on Large Vision-Language Models (LVLMs) and explores the policies of multi-modal demonstration selection. Existing research efforts in ICL face significant challenges: First, they rely on pre-defined demonstrations or heuristic selecting strategies based on human intuition, which are usually inadequate for covering diverse task requirements, leading to sub-optimal solutions; Second, individually selecting each demonstration fails in modeling the interactions between them, resulting in information redundancy. Unlike these prevailing efforts, we propose a new exploration-exploitation reinforcement learning framework, which explores policies to fuse multi-modal information and adaptively select adequate demonstrations as an integrated whole. The framework allows LVLMs to optimize themselves by continually refining their demonstrations through self-exploration, enabling the ability to autonomously identify and generate the most effective selection policies for in-context learning. Experimental results verify the superior performance of our approach on four Visual Question-Answering (VQA) datasets, demonstrating its effectiveness in enhancing the generalization capability of few-shot LVLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09473
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning
Chen, Cheng
Zhai, Yunpeng
Zhao, Yifan
Gao, Jinyang
Ding, Bolin
Li, Jia
Computer Vision and Pattern Recognition
In-context learning (ICL), a predominant trend in instruction learning, aims at enhancing the performance of large language models by providing clear task guidance and examples, improving their capability in task understanding and execution. This paper investigates ICL on Large Vision-Language Models (LVLMs) and explores the policies of multi-modal demonstration selection. Existing research efforts in ICL face significant challenges: First, they rely on pre-defined demonstrations or heuristic selecting strategies based on human intuition, which are usually inadequate for covering diverse task requirements, leading to sub-optimal solutions; Second, individually selecting each demonstration fails in modeling the interactions between them, resulting in information redundancy. Unlike these prevailing efforts, we propose a new exploration-exploitation reinforcement learning framework, which explores policies to fuse multi-modal information and adaptively select adequate demonstrations as an integrated whole. The framework allows LVLMs to optimize themselves by continually refining their demonstrations through self-exploration, enabling the ability to autonomously identify and generate the most effective selection policies for in-context learning. Experimental results verify the superior performance of our approach on four Visual Question-Answering (VQA) datasets, demonstrating its effectiveness in enhancing the generalization capability of few-shot LVLMs.
title Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.09473