Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Zijin, Likhomanenko, Tatiana, Jaitly, Navdeep
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918186753982464
author Gu, Zijin
Likhomanenko, Tatiana
Jaitly, Navdeep
author_facet Gu, Zijin
Likhomanenko, Tatiana
Jaitly, Navdeep
contents Mixture-of-experts (MoE) architectures have expanded from language modeling to automatic speech recognition (ASR). Traditional MoE methods, such as the Switch Transformer, route experts independently within each layer. Our analysis reveals that routers in most layers make expert choices that are not strongly correlated with the choices of the routers in other layers. To increase the cooperation between experts in different layers and encourage greater specialization, we use a shared router across different MoE layers. We call this model Omni-router Transformer. Extensive experiments on a large-scale pseudo-labeled dataset and evaluations across 10 diverse, out-of-domain ASR benchmarks demonstrate that the Omni-router Transformer is able to achieve lower training loss and consistently outperform dense and Switch Transformer models, reducing average word error rates by 11.2% and 8.2%, respectively, while providing structured expert usage and improved robustness to diverse data.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05724
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
Gu, Zijin
Likhomanenko, Tatiana
Jaitly, Navdeep
Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
Mixture-of-experts (MoE) architectures have expanded from language modeling to automatic speech recognition (ASR). Traditional MoE methods, such as the Switch Transformer, route experts independently within each layer. Our analysis reveals that routers in most layers make expert choices that are not strongly correlated with the choices of the routers in other layers. To increase the cooperation between experts in different layers and encourage greater specialization, we use a shared router across different MoE layers. We call this model Omni-router Transformer. Extensive experiments on a large-scale pseudo-labeled dataset and evaluations across 10 diverse, out-of-domain ASR benchmarks demonstrate that the Omni-router Transformer is able to achieve lower training loss and consistently outperform dense and Switch Transformer models, reducing average word error rates by 11.2% and 8.2%, respectively, while providing structured expert usage and improved robustness to diverse data.
title Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
topic Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2507.05724