Parrot: Multilingual Visual Instruction Tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Hai-Long, Zhou, Da-Wei, Li, Yang, Lu, Shiyin, Yi, Chao, Chen, Qing-Guo, Xu, Zhao, Luo, Weihua, Zhang, Kaifu, Zhan, De-Chuan, Ye, Han-Jia
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918034456707072
author Sun, Hai-Long
Zhou, Da-Wei
Li, Yang
Lu, Shiyin
Yi, Chao
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
Zhan, De-Chuan
Ye, Han-Jia
author_facet Sun, Hai-Long
Zhou, Da-Wei
Li, Yang
Lu, Shiyin
Yi, Chao
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
Zhan, De-Chuan
Ye, Han-Jia
contents The rapid development of Multimodal Large Language Models (MLLMs), such as GPT-4o, marks a significant step toward artificial general intelligence. Existing methods typically align vision encoders with LLMs via supervised fine-tuning (SFT), but this often deteriorates their ability to handle multiple languages as training progresses. We empirically observe that imbalanced SFT datasets, largely English-centric, degrade performance on non-English languages due to the failure in multilingual token alignment. To address this, we propose PARROT, a novel approach that leverages textual guidance for visual token alignment at the language level. PARROT conditions visual tokens on diverse language inputs and uses Mixture-of-Experts (MoE) to align multilingual tokens. By computing cross-attention between initial visual features and textual embeddings, we select the most relevant experts, converting visual tokens into language-specific representations. Additionally, we introduce the Massive Multilingual Multimodal Benchmark (MMMB), a new benchmark comprising 6 languages, 15 categories, and 12,000 questions, to assess multilingual capabilities. PARROT achieves state-of-the-art performance on both the multilingual benchmarks and a wide range of multimodal tasks. Code and dataset are available at: https://github.com/AIDC-AI/Parrot
format Preprint
id arxiv_https___arxiv_org_abs_2406_02539
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Parrot: Multilingual Visual Instruction Tuning
Sun, Hai-Long
Zhou, Da-Wei
Li, Yang
Lu, Shiyin
Yi, Chao
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
Zhan, De-Chuan
Ye, Han-Jia
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
The rapid development of Multimodal Large Language Models (MLLMs), such as GPT-4o, marks a significant step toward artificial general intelligence. Existing methods typically align vision encoders with LLMs via supervised fine-tuning (SFT), but this often deteriorates their ability to handle multiple languages as training progresses. We empirically observe that imbalanced SFT datasets, largely English-centric, degrade performance on non-English languages due to the failure in multilingual token alignment. To address this, we propose PARROT, a novel approach that leverages textual guidance for visual token alignment at the language level. PARROT conditions visual tokens on diverse language inputs and uses Mixture-of-Experts (MoE) to align multilingual tokens. By computing cross-attention between initial visual features and textual embeddings, we select the most relevant experts, converting visual tokens into language-specific representations. Additionally, we introduce the Massive Multilingual Multimodal Benchmark (MMMB), a new benchmark comprising 6 languages, 15 categories, and 12,000 questions, to assess multilingual capabilities. PARROT achieves state-of-the-art performance on both the multilingual benchmarks and a wide range of multimodal tasks. Code and dataset are available at: https://github.com/AIDC-AI/Parrot
title Parrot: Multilingual Visual Instruction Tuning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2406.02539