Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Geigle, Gregor, Schneider, Florian, Holtermann, Carolin, Biemann, Chris, Timofte, Radu, Lauscher, Anne, Glavaš, Goran
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929668112777216
author Geigle, Gregor
Schneider, Florian
Holtermann, Carolin
Biemann, Chris
Timofte, Radu
Lauscher, Anne
Glavaš, Goran
author_facet Geigle, Gregor
Schneider, Florian
Holtermann, Carolin
Biemann, Chris
Timofte, Radu
Lauscher, Anne
Glavaš, Goran
contents Most Large Vision-Language Models (LVLMs) to date are trained predominantly on English data, which makes them struggle to understand non-English input and fail to generate output in the desired target language. Existing efforts mitigate these issues by adding multilingual training data, but do so in a largely ad-hoc manner, lacking insight into how different training mixes tip the scale for different groups of languages. In this work, we present a comprehensive investigation into the training strategies for massively multilingual LVLMs. First, we conduct a series of multi-stage experiments spanning 13 downstream vision-language tasks and 43 languages, systematically examining: (1) the number of training languages that can be included without degrading English performance and (2) optimal language distributions of pre-training as well as (3) instruction-tuning data. Further, we (4) investigate how to improve multilingual text-in-image understanding, and introduce a new benchmark for the task. Surprisingly, our analysis reveals that one can (i) include as many as 100 training languages simultaneously (ii) with as little as 25-50\% of non-English data, to greatly improve multilingual performance while retaining strong English performance. We further find that (iii) including non-English OCR data in pre-training and instruction-tuning is paramount for improving multilingual text-in-image understanding. Finally, we put all our findings together and train Centurio, a 100-language LVLM, offering state-of-the-art performance in an evaluation covering 14 tasks and 56 languages.
format Preprint
id arxiv_https___arxiv_org_abs_2501_05122
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model
Geigle, Gregor
Schneider, Florian
Holtermann, Carolin
Biemann, Chris
Timofte, Radu
Lauscher, Anne
Glavaš, Goran
Computation and Language
Computer Vision and Pattern Recognition
Most Large Vision-Language Models (LVLMs) to date are trained predominantly on English data, which makes them struggle to understand non-English input and fail to generate output in the desired target language. Existing efforts mitigate these issues by adding multilingual training data, but do so in a largely ad-hoc manner, lacking insight into how different training mixes tip the scale for different groups of languages. In this work, we present a comprehensive investigation into the training strategies for massively multilingual LVLMs. First, we conduct a series of multi-stage experiments spanning 13 downstream vision-language tasks and 43 languages, systematically examining: (1) the number of training languages that can be included without degrading English performance and (2) optimal language distributions of pre-training as well as (3) instruction-tuning data. Further, we (4) investigate how to improve multilingual text-in-image understanding, and introduce a new benchmark for the task. Surprisingly, our analysis reveals that one can (i) include as many as 100 training languages simultaneously (ii) with as little as 25-50\% of non-English data, to greatly improve multilingual performance while retaining strong English performance. We further find that (iii) including non-English OCR data in pre-training and instruction-tuning is paramount for improving multilingual text-in-image understanding. Finally, we put all our findings together and train Centurio, a 100-language LVLM, offering state-of-the-art performance in an evaluation covering 14 tasks and 56 languages.
title Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.05122