Mimic In-Context Learning for Multimodal Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Yuchu, Fu, Jiale, Hao, Chenduo, Hu, Xinting, Peng, Yingzhe, Geng, Xin, Yang, Xu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909614190100480
author Jiang, Yuchu
Fu, Jiale
Hao, Chenduo
Hu, Xinting
Peng, Yingzhe
Geng, Xin
Yang, Xu
author_facet Jiang, Yuchu
Fu, Jiale
Hao, Chenduo
Hu, Xinting
Peng, Yingzhe
Geng, Xin
Yang, Xu
contents Recently, In-context Learning (ICL) has become a significant inference paradigm in Large Multimodal Models (LMMs), utilizing a few in-context demonstrations (ICDs) to prompt LMMs for new tasks. However, the synergistic effects in multimodal data increase the sensitivity of ICL performance to the configurations of ICDs, stimulating the need for a more stable and general mapping function. Mathematically, in Transformer-based models, ICDs act as "shift vectors" added to the hidden states of query tokens. Inspired by this, we introduce Mimic In-Context Learning (MimIC) to learn stable and generalizable shift effects from ICDs. Specifically, compared with some previous shift vector-based methods, MimIC more strictly approximates the shift effects by integrating lightweight learnable modules into LMMs with four key enhancements: 1) inserting shift vectors after attention layers, 2) assigning a shift vector to each attention head, 3) making shift magnitude query-dependent, and 4) employing a layer-wise alignment loss. Extensive experiments on two LMMs (Idefics-9b and Idefics2-8b-base) across three multimodal tasks (VQAv2, OK-VQA, Captioning) demonstrate that MimIC outperforms existing shift vector-based methods. The code is available at https://github.com/Kamichanw/MimIC.
format Preprint
id arxiv_https___arxiv_org_abs_2504_08851
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mimic In-Context Learning for Multimodal Tasks
Jiang, Yuchu
Fu, Jiale
Hao, Chenduo
Hu, Xinting
Peng, Yingzhe
Geng, Xin
Yang, Xu
Machine Learning
Artificial Intelligence
Recently, In-context Learning (ICL) has become a significant inference paradigm in Large Multimodal Models (LMMs), utilizing a few in-context demonstrations (ICDs) to prompt LMMs for new tasks. However, the synergistic effects in multimodal data increase the sensitivity of ICL performance to the configurations of ICDs, stimulating the need for a more stable and general mapping function. Mathematically, in Transformer-based models, ICDs act as "shift vectors" added to the hidden states of query tokens. Inspired by this, we introduce Mimic In-Context Learning (MimIC) to learn stable and generalizable shift effects from ICDs. Specifically, compared with some previous shift vector-based methods, MimIC more strictly approximates the shift effects by integrating lightweight learnable modules into LMMs with four key enhancements: 1) inserting shift vectors after attention layers, 2) assigning a shift vector to each attention head, 3) making shift magnitude query-dependent, and 4) employing a layer-wise alignment loss. Extensive experiments on two LMMs (Idefics-9b and Idefics2-8b-base) across three multimodal tasks (VQAv2, OK-VQA, Captioning) demonstrate that MimIC outperforms existing shift vector-based methods. The code is available at https://github.com/Kamichanw/MimIC.
title Mimic In-Context Learning for Multimodal Tasks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.08851