X-LLaVA: Optimizing Bilingual Large Vision-Language Alignment

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Shin, Dongjae, Lim, Hyeonseok, Won, Inho, Choi, Changsu, Kim, Minjun, Song, Seungwoo, Yoo, Hangyeol, Kim, Sangmin, Lim, Kyungtae
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929446270795776
author Shin, Dongjae
Lim, Hyeonseok
Won, Inho
Choi, Changsu
Kim, Minjun
Song, Seungwoo
Yoo, Hangyeol
Kim, Sangmin
Lim, Kyungtae
author_facet Shin, Dongjae
Lim, Hyeonseok
Won, Inho
Choi, Changsu
Kim, Minjun
Song, Seungwoo
Yoo, Hangyeol
Kim, Sangmin
Lim, Kyungtae
contents The impressive development of large language models (LLMs) is expanding into the realm of large multimodal models (LMMs), which incorporate multiple types of data beyond text. However, the nature of multimodal models leads to significant expenses in the creation of training data. Furthermore, constructing multilingual data for LMMs presents its own set of challenges due to language diversity and complexity. Therefore, in this study, we propose two cost-effective methods to solve this problem: (1) vocabulary expansion and pretraining of multilingual LLM for specific languages, and (2) automatic and elaborate construction of multimodal datasets using GPT4-V. Based on015 these methods, we constructed a 91K English-Korean-Chinese multilingual, multimodal training dataset. Additionally, we developed a bilingual multimodal model that exhibits excellent performance in both Korean and English, surpassing existing approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2403_11399
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle X-LLaVA: Optimizing Bilingual Large Vision-Language Alignment
Shin, Dongjae
Lim, Hyeonseok
Won, Inho
Choi, Changsu
Kim, Minjun
Song, Seungwoo
Yoo, Hangyeol
Kim, Sangmin
Lim, Kyungtae
Computation and Language
The impressive development of large language models (LLMs) is expanding into the realm of large multimodal models (LMMs), which incorporate multiple types of data beyond text. However, the nature of multimodal models leads to significant expenses in the creation of training data. Furthermore, constructing multilingual data for LMMs presents its own set of challenges due to language diversity and complexity. Therefore, in this study, we propose two cost-effective methods to solve this problem: (1) vocabulary expansion and pretraining of multilingual LLM for specific languages, and (2) automatic and elaborate construction of multimodal datasets using GPT4-V. Based on015 these methods, we constructed a 91K English-Korean-Chinese multilingual, multimodal training dataset. Additionally, we developed a bilingual multimodal model that exhibits excellent performance in both Korean and English, surpassing existing approaches.
title X-LLaVA: Optimizing Bilingual Large Vision-Language Alignment
topic Computation and Language
url https://arxiv.org/abs/2403.11399