Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Bo, Dong, Tianyu, Zhu, Shaolin, Xiong, Deyi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910272803831808
author Li, Bo
Dong, Tianyu
Zhu, Shaolin
Xiong, Deyi
author_facet Li, Bo
Dong, Tianyu
Zhu, Shaolin
Xiong, Deyi
contents Large Language Models (LLMs) have shown great promise in multilingual machine translation (MT), even with limited bilingual supervision. However, fine-tuning LLMs with parallel corpora presents major challenges, namely parameter interference. To address these issues, we propose Mix-MoE, a mixed Mixture-of-Experts framework designed to train LLMs for multilingual MT. Our framework operates in two distinct stages: (1) post-pretraining with MoE on monolingual corpora, and (2) post-pretraining with MoE on parallel corpora. Crucially, we divide the MoE layers into two specialized groups: Language Model Experts (LM Experts) and Machine Translation Experts (MT Experts). LM Experts are designed to capture and retain the monolingual knowledge learned by the pre-trained LLM. MT Experts, on the other hand, are specifically trained to acquire and store bilingual translation knowledge. Furthermore, to facilitate effective interaction between these specialized experts and leverage potential underlying structural patterns in text, we introduce a routing mechanism enhanced by Fourier Transform features derived from model representations. The experimental results demonstrate that Mix-MoE excels in multilingual MT, significantly outperforming existing baselines and showing notable progress in mitigating parameter interference.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24681
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs
Li, Bo
Dong, Tianyu
Zhu, Shaolin
Xiong, Deyi
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) have shown great promise in multilingual machine translation (MT), even with limited bilingual supervision. However, fine-tuning LLMs with parallel corpora presents major challenges, namely parameter interference. To address these issues, we propose Mix-MoE, a mixed Mixture-of-Experts framework designed to train LLMs for multilingual MT. Our framework operates in two distinct stages: (1) post-pretraining with MoE on monolingual corpora, and (2) post-pretraining with MoE on parallel corpora. Crucially, we divide the MoE layers into two specialized groups: Language Model Experts (LM Experts) and Machine Translation Experts (MT Experts). LM Experts are designed to capture and retain the monolingual knowledge learned by the pre-trained LLM. MT Experts, on the other hand, are specifically trained to acquire and store bilingual translation knowledge. Furthermore, to facilitate effective interaction between these specialized experts and leverage potential underlying structural patterns in text, we introduce a routing mechanism enhanced by Fourier Transform features derived from model representations. The experimental results demonstrate that Mix-MoE excels in multilingual MT, significantly outperforming existing baselines and showing notable progress in mitigating parameter interference.
title Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.24681