Multilingual Training and Evaluation Resources for Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baiamonte, Daniela, Fano, Elena, Gabburo, Matteo, Simonazzi, Stefano, Rigutini, Leonardo, Zugarini, Andrea
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910148938694656
author Baiamonte, Daniela
Fano, Elena
Gabburo, Matteo
Simonazzi, Stefano
Rigutini, Leonardo
Zugarini, Andrea
author_facet Baiamonte, Daniela
Fano, Elena
Gabburo, Matteo
Simonazzi, Stefano
Rigutini, Leonardo
Zugarini, Andrea
contents Vision Language Models (VLMs) achieved rapid progress in the recent years. However, despite their growth, VLMs development is heavily grounded on English, leading to two main limitations: (i) the lack of multilingual and multimodal datasets for training, and (ii) the scarcity of comprehensive evaluation benchmarks across languages. In this work, we address these gaps by introducing a new comprehensive suite of resources for VLMs training and evaluation spanning five European languages (English, French, German, Italian, and Spanish). We adopt a regeneration-translation paradigm that produces high-quality cross-lingual resources by combining curated synthetic generation and manual annotation. Specifically, we build Multi-PixMo, a training corpus obtained regenerating examples from Pixmo pre-existing datasets with permissively licensed models: PixMo-Cap, PixMo-AskModelAnything, and CoSyn-400k. On the evaluation side, we construct a set of multilingual benchmarks derived translating widely used English datasets (MMbench, ScienceQA, MME, POPE, AI2D). We assess the quality of these resources through qualitative and quantitative human analyses, measuring inter-annotator agreement. Additionally, we perform ablation studies to demonstrate the impact of multilingual data, with respect to English only, in VLMs training. Experiments, comprising 3 different models show that using multilingual, multimodal examples for training VLMs aids is consistently beneficial on non-English benchmarks, with positive transfer to English as well.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18347
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multilingual Training and Evaluation Resources for Vision-Language Models
Baiamonte, Daniela
Fano, Elena
Gabburo, Matteo
Simonazzi, Stefano
Rigutini, Leonardo
Zugarini, Andrea
Computation and Language
Artificial Intelligence
Vision Language Models (VLMs) achieved rapid progress in the recent years. However, despite their growth, VLMs development is heavily grounded on English, leading to two main limitations: (i) the lack of multilingual and multimodal datasets for training, and (ii) the scarcity of comprehensive evaluation benchmarks across languages. In this work, we address these gaps by introducing a new comprehensive suite of resources for VLMs training and evaluation spanning five European languages (English, French, German, Italian, and Spanish). We adopt a regeneration-translation paradigm that produces high-quality cross-lingual resources by combining curated synthetic generation and manual annotation. Specifically, we build Multi-PixMo, a training corpus obtained regenerating examples from Pixmo pre-existing datasets with permissively licensed models: PixMo-Cap, PixMo-AskModelAnything, and CoSyn-400k. On the evaluation side, we construct a set of multilingual benchmarks derived translating widely used English datasets (MMbench, ScienceQA, MME, POPE, AI2D). We assess the quality of these resources through qualitative and quantitative human analyses, measuring inter-annotator agreement. Additionally, we perform ablation studies to demonstrate the impact of multilingual data, with respect to English only, in VLMs training. Experiments, comprising 3 different models show that using multilingual, multimodal examples for training VLMs aids is consistently beneficial on non-English benchmarks, with positive transfer to English as well.
title Multilingual Training and Evaluation Resources for Vision-Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.18347