What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Qin, Libo, Chen, Qiguang, Fei, Hao, Chen, Zhi, Li, Min, Che, Wanxiang
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916456734654464
author Qin, Libo
Chen, Qiguang
Fei, Hao
Chen, Zhi
Li, Min
Che, Wanxiang
author_facet Qin, Libo
Chen, Qiguang
Fei, Hao
Chen, Zhi
Li, Min
Che, Wanxiang
contents Recently, rapid advancements in Multi-Modal In-Context Learning (MM-ICL) have achieved notable success, which is capable of achieving superior performance across various tasks without requiring additional parameter tuning. However, the underlying rules for the effectiveness of MM-ICL remain under-explored. To fill this gap, this work aims to investigate the research question: "What factors affect the performance of MM-ICL?'' To this end, we investigate extensive experiments on the three core steps of MM-ICL including demonstration retrieval, demonstration ordering, and prompt construction using 6 vision large language models and 20 strategies. Our findings highlight (1) the necessity of a multi-modal retriever for demonstration retrieval, (2) the importance of intra-demonstration ordering over inter-demonstration ordering, and (3) the enhancement of task comprehension through introductory instructions in prompts. We hope this study can serve as a foundational guide for optimizing MM-ICL strategies in future research.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20482
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
Qin, Libo
Chen, Qiguang
Fei, Hao
Chen, Zhi
Li, Min
Che, Wanxiang
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Recently, rapid advancements in Multi-Modal In-Context Learning (MM-ICL) have achieved notable success, which is capable of achieving superior performance across various tasks without requiring additional parameter tuning. However, the underlying rules for the effectiveness of MM-ICL remain under-explored. To fill this gap, this work aims to investigate the research question: "What factors affect the performance of MM-ICL?'' To this end, we investigate extensive experiments on the three core steps of MM-ICL including demonstration retrieval, demonstration ordering, and prompt construction using 6 vision large language models and 20 strategies. Our findings highlight (1) the necessity of a multi-modal retriever for demonstration retrieval, (2) the importance of intra-demonstration ordering over inter-demonstration ordering, and (3) the enhancement of task comprehension through introductory instructions in prompts. We hope this study can serve as a foundational guide for optimizing MM-ICL strategies in future research.
title What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.20482