Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Shuo, Han, Zhen, He, Bailan, Liu, Jianzhe, Buckley, Mark, Qin, Yao, Torr, Philip, Tresp, Volker, Gu, Jindong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916511497584640
author Chen, Shuo
Han, Zhen
He, Bailan
Liu, Jianzhe
Buckley, Mark
Qin, Yao
Torr, Philip
Tresp, Volker
Gu, Jindong
author_facet Chen, Shuo
Han, Zhen
He, Bailan
Liu, Jianzhe
Buckley, Mark
Qin, Yao
Torr, Philip
Tresp, Volker
Gu, Jindong
contents Large Language Models (LLMs) with in-context learning (ICL) ability can quickly adapt to a specific context given a few demonstrations (demos). Recently, Multimodal Large Language Models (MLLMs) built upon LLMs have also shown multimodal ICL ability, i.e., responding to queries given a few multimodal demos, including images, queries, and answers. While ICL has been extensively studied on LLMs, its research on MLLMs remains limited. One essential question is whether these MLLMs can truly conduct multimodal ICL, or if only the textual modality is necessary. We investigate this question by examining two primary factors that influence ICL: 1) Demo content, i.e., understanding the influences of demo content in different modalities. 2) Demo selection strategy, i.e., how to select better multimodal demos for improved performance. Experiments revealed that multimodal ICL is predominantly driven by the textual content whereas the visual information in the demos has little influence. Interestingly, visual content is still necessary and useful for selecting demos to increase performance. Motivated by our analysis, we propose a simple yet effective approach, termed Mixed Modality In-Context Example Selection (MMICES), which considers both visual and language modalities when selecting demos. Extensive experiments are conducted to support our findings and verify the improvement brought by our method. Code is available at \url{https://chenxshuo.github.io/m-icl/}.
format Preprint
id arxiv_https___arxiv_org_abs_2311_18021
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?
Chen, Shuo
Han, Zhen
He, Bailan
Liu, Jianzhe
Buckley, Mark
Qin, Yao
Torr, Philip
Tresp, Volker
Gu, Jindong
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Large Language Models (LLMs) with in-context learning (ICL) ability can quickly adapt to a specific context given a few demonstrations (demos). Recently, Multimodal Large Language Models (MLLMs) built upon LLMs have also shown multimodal ICL ability, i.e., responding to queries given a few multimodal demos, including images, queries, and answers. While ICL has been extensively studied on LLMs, its research on MLLMs remains limited. One essential question is whether these MLLMs can truly conduct multimodal ICL, or if only the textual modality is necessary. We investigate this question by examining two primary factors that influence ICL: 1) Demo content, i.e., understanding the influences of demo content in different modalities. 2) Demo selection strategy, i.e., how to select better multimodal demos for improved performance. Experiments revealed that multimodal ICL is predominantly driven by the textual content whereas the visual information in the demos has little influence. Interestingly, visual content is still necessary and useful for selecting demos to increase performance. Motivated by our analysis, we propose a simple yet effective approach, termed Mixed Modality In-Context Example Selection (MMICES), which considers both visual and language modalities when selecting demos. Extensive experiments are conducted to support our findings and verify the improvement brought by our method. Code is available at \url{https://chenxshuo.github.io/m-icl/}.
title Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2311.18021