MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Han, Wenhan, Zhang, Yifan, Chen, Zhixun, Liu, Binbin, Lin, Haobin, Zhang, Bingni, Wang, Taifeng, Pechenizkiy, Mykola, Fang, Meng, Zheng, Yin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915357571153920
author Han, Wenhan
Zhang, Yifan
Chen, Zhixun
Liu, Binbin
Lin, Haobin
Zhang, Bingni
Wang, Taifeng
Pechenizkiy, Mykola
Fang, Meng
Zheng, Yin
author_facet Han, Wenhan
Zhang, Yifan
Chen, Zhixun
Liu, Binbin
Lin, Haobin
Zhang, Bingni
Wang, Taifeng
Pechenizkiy, Mykola
Fang, Meng
Zheng, Yin
contents Multilingual large language models (LLMs) are advancing rapidly, with new models frequently claiming support for an increasing number of languages. However, existing evaluation datasets are limited and lack cross-lingual alignment, leaving assessments of multilingual capabilities fragmented in both language and skill coverage. To address this, we introduce MuBench, a benchmark covering 61 languages and evaluating a broad range of capabilities. We evaluate several state-of-the-art multilingual LLMs and find notable gaps between claimed and actual language coverage, particularly a persistent performance disparity between English and low-resource languages. Leveraging MuBench's alignment, we propose Multilingual Consistency (MLC) as a complementary metric to accuracy for analyzing performance bottlenecks and guiding model improvement. Finally, we pretrain a suite of 1.2B-parameter models on English and Chinese with 500B tokens, varying language ratios and parallel data proportions to investigate cross-lingual transfer dynamics.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19468
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
Han, Wenhan
Zhang, Yifan
Chen, Zhixun
Liu, Binbin
Lin, Haobin
Zhang, Bingni
Wang, Taifeng
Pechenizkiy, Mykola
Fang, Meng
Zheng, Yin
Computation and Language
Artificial Intelligence
Multilingual large language models (LLMs) are advancing rapidly, with new models frequently claiming support for an increasing number of languages. However, existing evaluation datasets are limited and lack cross-lingual alignment, leaving assessments of multilingual capabilities fragmented in both language and skill coverage. To address this, we introduce MuBench, a benchmark covering 61 languages and evaluating a broad range of capabilities. We evaluate several state-of-the-art multilingual LLMs and find notable gaps between claimed and actual language coverage, particularly a persistent performance disparity between English and low-resource languages. Leveraging MuBench's alignment, we propose Multilingual Consistency (MLC) as a complementary metric to accuracy for analyzing performance bottlenecks and guiding model improvement. Finally, we pretrain a suite of 1.2B-parameter models on English and Chinese with 500B tokens, varying language ratios and parallel data proportions to investigate cross-lingual transfer dynamics.
title MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.19468