Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Xiujun, Lu, Yujie, Gan, Zhe, Gao, Jianfeng, Wang, William Yang, Choi, Yejin
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917689583206400
author Li, Xiujun
Lu, Yujie
Gan, Zhe
Gao, Jianfeng
Wang, William Yang
Choi, Yejin
author_facet Li, Xiujun
Lu, Yujie
Gan, Zhe
Gao, Jianfeng
Wang, William Yang
Choi, Yejin
contents Recent multimodal large language models (MLLMs) have shown promising instruction following capabilities on vision-language tasks. In this work, we introduce VISUAL MODALITY INSTRUCTION (VIM), and investigate how well multimodal models can understand textual instructions provided in pixels, despite not being explicitly trained on such data during pretraining or fine-tuning. We adapt VIM to eight benchmarks, including OKVQA, MM-Vet, MathVista, MMMU, and probe diverse MLLMs in both the text-modality instruction (TEM) setting and VIM setting. Notably, we observe a significant performance disparity between the original TEM and VIM settings for open-source MLLMs, indicating that open-source MLLMs face greater challenges when text instruction is presented solely in image form. To address this issue, we train v-MLLM, a generalizable model that is capable to conduct robust instruction following in both text-modality and visual-modality instructions.
format Preprint
id arxiv_https___arxiv_org_abs_2311_17647
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
Li, Xiujun
Lu, Yujie
Gan, Zhe
Gao, Jianfeng
Wang, William Yang
Choi, Yejin
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Recent multimodal large language models (MLLMs) have shown promising instruction following capabilities on vision-language tasks. In this work, we introduce VISUAL MODALITY INSTRUCTION (VIM), and investigate how well multimodal models can understand textual instructions provided in pixels, despite not being explicitly trained on such data during pretraining or fine-tuning. We adapt VIM to eight benchmarks, including OKVQA, MM-Vet, MathVista, MMMU, and probe diverse MLLMs in both the text-modality instruction (TEM) setting and VIM setting. Notably, we observe a significant performance disparity between the original TEM and VIM settings for open-source MLLMs, indicating that open-source MLLMs face greater challenges when text instruction is presented solely in image form. To address this issue, we train v-MLLM, a generalizable model that is capable to conduct robust instruction following in both text-modality and visual-modality instructions.
title Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2311.17647