Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Chong, Deng, Yingzhuo, Zhang, Jiajun, Zong, Chengqing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911006490361856
author Li, Chong
Deng, Yingzhuo
Zhang, Jiajun
Zong, Chengqing
author_facet Li, Chong
Deng, Yingzhuo
Zhang, Jiajun
Zong, Chengqing
contents The curse of multilinguality phenomenon is a fundamental problem of multilingual Large Language Models (LLMs), where the competition between massive languages results in inferior performance. It mainly comes from limited capacity and negative transfer between dissimilar languages. To address this issue, we propose a method to dynamically group and scale up the parameters of multilingual LLM while boosting positive transfer among similar languages. Specifically, the model is first tuned on monolingual corpus to determine the parameter deviation in each layer and quantify the similarity between languages. Layers with more deviations are extended to mixture-of-experts layers to reduce competition between languages, where one expert module serves one group of similar languages. Experimental results on 18 to 128 languages show that our method reduces the negative transfer between languages and significantly boosts multilingual performance with fewer parameters. Such language group specialization on experts benefits the new language adaptation and reduces the inference on the previous multilingual knowledge learned.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12388
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
Li, Chong
Deng, Yingzhuo
Zhang, Jiajun
Zong, Chengqing
Computation and Language
Artificial Intelligence
The curse of multilinguality phenomenon is a fundamental problem of multilingual Large Language Models (LLMs), where the competition between massive languages results in inferior performance. It mainly comes from limited capacity and negative transfer between dissimilar languages. To address this issue, we propose a method to dynamically group and scale up the parameters of multilingual LLM while boosting positive transfer among similar languages. Specifically, the model is first tuned on monolingual corpus to determine the parameter deviation in each layer and quantify the similarity between languages. Layers with more deviations are extended to mixture-of-experts layers to reduce competition between languages, where one expert module serves one group of similar languages. Experimental results on 18 to 128 languages show that our method reduces the negative transfer between languages and significantly boosts multilingual performance with fewer parameters. Such language group specialization on experts benefits the new language adaptation and reduces the inference on the previous multilingual knowledge learned.
title Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.12388