MMICT: Boosting Multi-Modal Fine-Tuning with In-Context Examples

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Tao, Zhang, Enwei, Gao, Yuting, Li, Ke, Sun, Xing, Zhang, Yan, Li, Hui, Ji, Rongrong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916351808897024
author Chen, Tao
Zhang, Enwei
Gao, Yuting
Li, Ke
Sun, Xing
Zhang, Yan
Li, Hui
Ji, Rongrong
author_facet Chen, Tao
Zhang, Enwei
Gao, Yuting
Li, Ke
Sun, Xing
Zhang, Yan
Li, Hui
Ji, Rongrong
contents Although In-Context Learning (ICL) brings remarkable performance gains to Large Language Models (LLMs), the improvements remain lower than fine-tuning on downstream tasks. This paper introduces Multi-Modal In-Context Tuning (MMICT), a novel multi-modal fine-tuning paradigm that boosts multi-modal fine-tuning by fully leveraging the promising ICL capability of multi-modal LLMs (MM-LLMs). We propose the Multi-Modal Hub (M-Hub), a unified module that captures various multi-modal features according to different inputs and objectives. Based on M-Hub, MMICT enables MM-LLMs to learn from in-context visual-guided textual features and subsequently generate outputs conditioned on the textual-guided visual features. Moreover, leveraging the flexibility of M-Hub, we design a variety of in-context demonstrations. Extensive experiments on a diverse range of downstream multi-modal tasks demonstrate that MMICT significantly outperforms traditional fine-tuning strategy and the vanilla ICT method that directly takes the concatenation of all information from different modalities as input. Our implementation is available at: https://github.com/KDEGroup/MMICT.
format Preprint
id arxiv_https___arxiv_org_abs_2312_06363
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MMICT: Boosting Multi-Modal Fine-Tuning with In-Context Examples
Chen, Tao
Zhang, Enwei
Gao, Yuting
Li, Ke
Sun, Xing
Zhang, Yan
Li, Hui
Ji, Rongrong
Artificial Intelligence
Computation and Language
Machine Learning
Although In-Context Learning (ICL) brings remarkable performance gains to Large Language Models (LLMs), the improvements remain lower than fine-tuning on downstream tasks. This paper introduces Multi-Modal In-Context Tuning (MMICT), a novel multi-modal fine-tuning paradigm that boosts multi-modal fine-tuning by fully leveraging the promising ICL capability of multi-modal LLMs (MM-LLMs). We propose the Multi-Modal Hub (M-Hub), a unified module that captures various multi-modal features according to different inputs and objectives. Based on M-Hub, MMICT enables MM-LLMs to learn from in-context visual-guided textual features and subsequently generate outputs conditioned on the textual-guided visual features. Moreover, leveraging the flexibility of M-Hub, we design a variety of in-context demonstrations. Extensive experiments on a diverse range of downstream multi-modal tasks demonstrate that MMICT significantly outperforms traditional fine-tuning strategy and the vanilla ICT method that directly takes the concatenation of all information from different modalities as input. Our implementation is available at: https://github.com/KDEGroup/MMICT.
title MMICT: Boosting Multi-Modal Fine-Tuning with In-Context Examples
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2312.06363