MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915357571153920 |
|---|---|
| author | Han, Wenhan Zhang, Yifan Chen, Zhixun Liu, Binbin Lin, Haobin Zhang, Bingni Wang, Taifeng Pechenizkiy, Mykola Fang, Meng Zheng, Yin |
| author_facet | Han, Wenhan Zhang, Yifan Chen, Zhixun Liu, Binbin Lin, Haobin Zhang, Bingni Wang, Taifeng Pechenizkiy, Mykola Fang, Meng Zheng, Yin |
| contents | Multilingual large language models (LLMs) are advancing rapidly, with new models frequently claiming support for an increasing number of languages. However, existing evaluation datasets are limited and lack cross-lingual alignment, leaving assessments of multilingual capabilities fragmented in both language and skill coverage. To address this, we introduce MuBench, a benchmark covering 61 languages and evaluating a broad range of capabilities. We evaluate several state-of-the-art multilingual LLMs and find notable gaps between claimed and actual language coverage, particularly a persistent performance disparity between English and low-resource languages. Leveraging MuBench's alignment, we propose Multilingual Consistency (MLC) as a complementary metric to accuracy for analyzing performance bottlenecks and guiding model improvement. Finally, we pretrain a suite of 1.2B-parameter models on English and Chinese with 500B tokens, varying language ratios and parallel data proportions to investigate cross-lingual transfer dynamics. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_19468 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Han, Wenhan Zhang, Yifan Chen, Zhixun Liu, Binbin Lin, Haobin Zhang, Bingni Wang, Taifeng Pechenizkiy, Mykola Fang, Meng Zheng, Yin Computation and Language Artificial Intelligence Multilingual large language models (LLMs) are advancing rapidly, with new models frequently claiming support for an increasing number of languages. However, existing evaluation datasets are limited and lack cross-lingual alignment, leaving assessments of multilingual capabilities fragmented in both language and skill coverage. To address this, we introduce MuBench, a benchmark covering 61 languages and evaluating a broad range of capabilities. We evaluate several state-of-the-art multilingual LLMs and find notable gaps between claimed and actual language coverage, particularly a persistent performance disparity between English and low-resource languages. Leveraging MuBench's alignment, we propose Multilingual Consistency (MLC) as a complementary metric to accuracy for analyzing performance bottlenecks and guiding model improvement. Finally, we pretrain a suite of 1.2B-parameter models on English and Chinese with 500B tokens, varying language ratios and parallel data proportions to investigate cross-lingual transfer dynamics. |
| title | MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2506.19468 |