MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tu, Jianhong, Ni, Zhuohao, Crispino, Nicholas, Yu, Zihao, Bendersky, Michael, Gunel, Beliz, Jia, Ruoxi, Liu, Xin, Lyu, Lingjuan, Song, Dawn, Wang, Chenguang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913915826339840
author Tu, Jianhong
Ni, Zhuohao
Crispino, Nicholas
Yu, Zihao
Bendersky, Michael
Gunel, Beliz
Jia, Ruoxi
Liu, Xin
Lyu, Lingjuan
Song, Dawn
Wang, Chenguang
author_facet Tu, Jianhong
Ni, Zhuohao
Crispino, Nicholas
Yu, Zihao
Bendersky, Michael
Gunel, Beliz
Jia, Ruoxi
Liu, Xin
Lyu, Lingjuan
Song, Dawn
Wang, Chenguang
contents We present a novel visual instruction tuning strategy to improve the zero-shot task generalization of multimodal large language models by building a firm text-only knowledge base. Existing work lacks sufficient experimentation on the importance of each modality in the instruction tuning stage, often using a majority of vision-language data while keeping text-only data limited and fixing mixtures of modalities. By incorporating diverse text-only data in the visual instruction tuning stage, we vary vision-language data in various controlled experiments to investigate the importance of modality in visual instruction tuning. Our comprehensive evaluation shows that the text-heavy instruction tuning approach is able to perform on-par with traditional vision-heavy mixtures on both modalities across 12 general datasets while using as low as half the total training tokens. We find that simply increasing sufficiently diverse text-only data enables transfer of instruction following ability and domain knowledge across modalities while being more efficient than the vision-language approach.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10557
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models
Tu, Jianhong
Ni, Zhuohao
Crispino, Nicholas
Yu, Zihao
Bendersky, Michael
Gunel, Beliz
Jia, Ruoxi
Liu, Xin
Lyu, Lingjuan
Song, Dawn
Wang, Chenguang
Computation and Language
We present a novel visual instruction tuning strategy to improve the zero-shot task generalization of multimodal large language models by building a firm text-only knowledge base. Existing work lacks sufficient experimentation on the importance of each modality in the instruction tuning stage, often using a majority of vision-language data while keeping text-only data limited and fixing mixtures of modalities. By incorporating diverse text-only data in the visual instruction tuning stage, we vary vision-language data in various controlled experiments to investigate the importance of modality in visual instruction tuning. Our comprehensive evaluation shows that the text-heavy instruction tuning approach is able to perform on-par with traditional vision-heavy mixtures on both modalities across 12 general datasets while using as low as half the total training tokens. We find that simply increasing sufficiently diverse text-only data enables transfer of instruction following ability and domain knowledge across modalities while being more efficient than the vision-language approach.
title MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models
topic Computation and Language
url https://arxiv.org/abs/2411.10557