Multilingual Routing in Mixture-of-Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bandarkar, Lucas, Yang, Chenyuan, Fayyaz, Mohsen, Hu, Junlin, Peng, Nanyun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912910942404608
author Bandarkar, Lucas
Yang, Chenyuan
Fayyaz, Mohsen
Hu, Junlin
Peng, Nanyun
author_facet Bandarkar, Lucas
Yang, Chenyuan
Fayyaz, Mohsen
Hu, Junlin
Peng, Nanyun
contents Mixture-of-Experts (MoE) architectures have become the key to scaling modern LLMs, yet little is understood about how their sparse routing dynamics respond to multilingual data. In this work, we analyze expert routing patterns using parallel multilingual datasets and present highly interpretable layer-wise phenomena. We find that MoE models route tokens in language-specific ways in the early and late decoder layers but exhibit significant cross-lingual routing alignment in middle layers, mirroring parameter-sharing trends observed in dense LLMs. In particular, we reveal a clear, strong correlation between a model's performance in a given language and how similarly its tokens are routed to English in these layers. Extending beyond correlation, we explore inference-time interventions that induce higher cross-lingual routing alignment. We introduce a method that steers the router by promoting middle-layer task experts frequently activated in English, and it successfully increases multilingual performance. These 1-2% gains are remarkably consistent across two evaluation tasks, three models, and 15+ languages, especially given that these simple interventions override routers of extensively trained, state-of-the-art LLMs. In comparison, interventions outside of the middle layers or targeting multilingual-specialized experts only yield performance degradation. Altogether, we present numerous findings that explain how MoEs process non-English text and demonstrate that generalization is limited by the model's ability to leverage language-universal experts in all languages.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04694
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multilingual Routing in Mixture-of-Experts
Bandarkar, Lucas
Yang, Chenyuan
Fayyaz, Mohsen
Hu, Junlin
Peng, Nanyun
Computation and Language
Artificial Intelligence
Machine Learning
Mixture-of-Experts (MoE) architectures have become the key to scaling modern LLMs, yet little is understood about how their sparse routing dynamics respond to multilingual data. In this work, we analyze expert routing patterns using parallel multilingual datasets and present highly interpretable layer-wise phenomena. We find that MoE models route tokens in language-specific ways in the early and late decoder layers but exhibit significant cross-lingual routing alignment in middle layers, mirroring parameter-sharing trends observed in dense LLMs. In particular, we reveal a clear, strong correlation between a model's performance in a given language and how similarly its tokens are routed to English in these layers. Extending beyond correlation, we explore inference-time interventions that induce higher cross-lingual routing alignment. We introduce a method that steers the router by promoting middle-layer task experts frequently activated in English, and it successfully increases multilingual performance. These 1-2% gains are remarkably consistent across two evaluation tasks, three models, and 15+ languages, especially given that these simple interventions override routers of extensively trained, state-of-the-art LLMs. In comparison, interventions outside of the middle layers or targeting multilingual-specialized experts only yield performance degradation. Altogether, we present numerous findings that explain how MoEs process non-English text and demonstrate that generalization is limited by the model's ability to leverage language-universal experts in all languages.
title Multilingual Routing in Mixture-of-Experts
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.04694