TowerVision: Understanding and Improving Multilinguality in Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Viveiros, André G., Fernandes, Patrick, Santos, Saul, Sannigrahi, Sonal, Zaranis, Emmanouil, Guerreiro, Nuno M., Farajian, Amin, Colombo, Pierre, Neubig, Graham, Martins, André F. T.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917063959773184
author Viveiros, André G.
Fernandes, Patrick
Santos, Saul
Sannigrahi, Sonal
Zaranis, Emmanouil
Guerreiro, Nuno M.
Farajian, Amin
Colombo, Pierre
Neubig, Graham
Martins, André F. T.
author_facet Viveiros, André G.
Fernandes, Patrick
Santos, Saul
Sannigrahi, Sonal
Zaranis, Emmanouil
Guerreiro, Nuno M.
Farajian, Amin
Colombo, Pierre
Neubig, Graham
Martins, André F. T.
contents Despite significant advances in vision-language models (VLMs), most existing work follows an English-centric design process, limiting their effectiveness in multilingual settings. In this work, we provide a comprehensive empirical study analyzing the impact of several multilingual design choices, such as training data composition, encoder selection, and text backbones. The result is TowerVision, a family of open multilingual VLMs for both image-text and video-text tasks, built upon the multilingual text-only model Tower+. TowerVision achieves competitive performance on multiple multimodal multilingual benchmarks and shows particular strength in culturally grounded tasks and multimodal translation. By incorporating visual and cultural context during fine-tuning, our models surpass existing approaches trained on substantially larger datasets, as demonstrated on ALM-Bench and Multi30K (image tasks) and ViMUL-Bench (video tasks). Alongside the models, we release VisionBlocks, a high-quality, curated vision-language dataset. Our findings highlight that multilingual vision-language training data substantially improves cross-lingual generalization -- both from high-resource to underrepresented languages and vice versa -- and that instruction-tuned LLMs are not always the optimal initialization point. To support further research, we publicly release all models, data, and training recipes.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21849
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
Viveiros, André G.
Fernandes, Patrick
Santos, Saul
Sannigrahi, Sonal
Zaranis, Emmanouil
Guerreiro, Nuno M.
Farajian, Amin
Colombo, Pierre
Neubig, Graham
Martins, André F. T.
Machine Learning
Artificial Intelligence
68T07, 68T45, 68T50
I.2.7; I.2.10; I.5.4
Despite significant advances in vision-language models (VLMs), most existing work follows an English-centric design process, limiting their effectiveness in multilingual settings. In this work, we provide a comprehensive empirical study analyzing the impact of several multilingual design choices, such as training data composition, encoder selection, and text backbones. The result is TowerVision, a family of open multilingual VLMs for both image-text and video-text tasks, built upon the multilingual text-only model Tower+. TowerVision achieves competitive performance on multiple multimodal multilingual benchmarks and shows particular strength in culturally grounded tasks and multimodal translation. By incorporating visual and cultural context during fine-tuning, our models surpass existing approaches trained on substantially larger datasets, as demonstrated on ALM-Bench and Multi30K (image tasks) and ViMUL-Bench (video tasks). Alongside the models, we release VisionBlocks, a high-quality, curated vision-language dataset. Our findings highlight that multilingual vision-language training data substantially improves cross-lingual generalization -- both from high-resource to underrepresented languages and vice versa -- and that instruction-tuned LLMs are not always the optimal initialization point. To support further research, we publicly release all models, data, and training recipes.
title TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
topic Machine Learning
Artificial Intelligence
68T07, 68T45, 68T50
I.2.7; I.2.10; I.5.4
url https://arxiv.org/abs/2510.21849