What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866916456734654464 |
|---|---|
| author | Qin, Libo Chen, Qiguang Fei, Hao Chen, Zhi Li, Min Che, Wanxiang |
| author_facet | Qin, Libo Chen, Qiguang Fei, Hao Chen, Zhi Li, Min Che, Wanxiang |
| contents | Recently, rapid advancements in Multi-Modal In-Context Learning (MM-ICL) have achieved notable success, which is capable of achieving superior performance across various tasks without requiring additional parameter tuning. However, the underlying rules for the effectiveness of MM-ICL remain under-explored. To fill this gap, this work aims to investigate the research question: "What factors affect the performance of MM-ICL?'' To this end, we investigate extensive experiments on the three core steps of MM-ICL including demonstration retrieval, demonstration ordering, and prompt construction using 6 vision large language models and 20 strategies. Our findings highlight (1) the necessity of a multi-modal retriever for demonstration retrieval, (2) the importance of intra-demonstration ordering over inter-demonstration ordering, and (3) the enhancement of task comprehension through introductory instructions in prompts. We hope this study can serve as a foundational guide for optimizing MM-ICL strategies in future research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_20482 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration Qin, Libo Chen, Qiguang Fei, Hao Chen, Zhi Li, Min Che, Wanxiang Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition Recently, rapid advancements in Multi-Modal In-Context Learning (MM-ICL) have achieved notable success, which is capable of achieving superior performance across various tasks without requiring additional parameter tuning. However, the underlying rules for the effectiveness of MM-ICL remain under-explored. To fill this gap, this work aims to investigate the research question: "What factors affect the performance of MM-ICL?'' To this end, we investigate extensive experiments on the three core steps of MM-ICL including demonstration retrieval, demonstration ordering, and prompt construction using 6 vision large language models and 20 strategies. Our findings highlight (1) the necessity of a multi-modal retriever for demonstration retrieval, (2) the importance of intra-demonstration ordering over inter-demonstration ordering, and (3) the enhancement of task comprehension through introductory instructions in prompts. We hope this study can serve as a foundational guide for optimizing MM-ICL strategies in future research. |
| title | What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration |
| topic | Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2410.20482 |