MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wei, Lai, Li, Xiaozhe, Jiang, Zihao, Huang, Weiran, Sun, Lichao
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911585337868288
author Wei, Lai
Li, Xiaozhe
Jiang, Zihao
Huang, Weiran
Sun, Lichao
author_facet Wei, Lai
Li, Xiaozhe
Jiang, Zihao
Huang, Weiran
Sun, Lichao
contents Multimodal large language models are typically trained in two stages: first pre-training on image-text pairs, and then fine-tuning using supervised vision-language instruction data. Recent studies have shown that large language models can achieve satisfactory results even with a limited amount of high-quality instruction-following data. In this paper, we introduce MM-LIMA, which is fine-tuned on a small dataset comprising only 200 examples, amounting to approximately 6% of the instruction-following data used in the alignment dataset for MiniGPT-4. To achieve this, we first propose several metrics to access the quality of multimodal instruction data. Based on these metrics, we present an effective and trainable data selector to automatically identify and filter low-quality vision-language data. By employing this method, MM-LIMA outperforms the original MiniGPT-4 on various evaluations. Overall, our findings demonstrate that less but high-quality instruction tuning data is efficient in enabling multimodal large language models to generate better output. Our code is available at https://github.com/waltonfuture/InstructionGPT-4.
format Preprint
id arxiv_https___arxiv_org_abs_2308_12067
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
Wei, Lai
Li, Xiaozhe
Jiang, Zihao
Huang, Weiran
Sun, Lichao
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Multimodal large language models are typically trained in two stages: first pre-training on image-text pairs, and then fine-tuning using supervised vision-language instruction data. Recent studies have shown that large language models can achieve satisfactory results even with a limited amount of high-quality instruction-following data. In this paper, we introduce MM-LIMA, which is fine-tuned on a small dataset comprising only 200 examples, amounting to approximately 6% of the instruction-following data used in the alignment dataset for MiniGPT-4. To achieve this, we first propose several metrics to access the quality of multimodal instruction data. Based on these metrics, we present an effective and trainable data selector to automatically identify and filter low-quality vision-language data. By employing this method, MM-LIMA outperforms the original MiniGPT-4 on various evaluations. Overall, our findings demonstrate that less but high-quality instruction tuning data is efficient in enabling multimodal large language models to generate better output. Our code is available at https://github.com/waltonfuture/InstructionGPT-4.
title MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2308.12067