Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhan, Zheng, Ren, Liliang, Wang, Shuohang, Liu, Liyuan, Liu, Yang, Gong, Yeyun, Wang, Yanzhi, Shen, Yelong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916806070894592
author Zhan, Zheng
Ren, Liliang
Wang, Shuohang
Liu, Liyuan
Liu, Yang
Gong, Yeyun
Wang, Yanzhi
Shen, Yelong
author_facet Zhan, Zheng
Ren, Liliang
Wang, Shuohang
Liu, Liyuan
Liu, Yang
Gong, Yeyun
Wang, Yanzhi
Shen, Yelong
contents Linear State Space Models (SSMs) offer remarkable performance gains in efficient sequence modeling, with constant inference-time computation and memory complexity. Recent advances, such as Mamba, further enhance SSMs with input-dependent gating and hardware-aware implementations, positioning them as strong alternatives to Transformers for long sequence modeling. However, efficiently scaling the expressive power of SSMs, particularly with Mixture of Experts (MoE), remains challenging, as naive integration attempts often falter or degrade performance. In this work, we introduce Routing Mamba (RoM), a novel approach that scales SSM parameters using sparse mixtures of linear projection experts. By sharing routing decisions between projection layers and lightweight sub-modules within Mamba across experts, RoM leverages synergies among linear projection experts for effective and efficient sparse scaling of Mamba layers. At a scale of 1.3B active parameters (10B total) and 16K training sequence length, RoM achieves language modeling performance equivalent to a dense Mamba model requiring over 2.3x more active parameters, and demonstrates consistent perplexity across context lengths. Experimental results further show RoM effectively scales hybrid language models, yielding a 23% FLOPS saving compared to dense Mamba scaling for similar performance.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18145
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
Zhan, Zheng
Ren, Liliang
Wang, Shuohang
Liu, Liyuan
Liu, Yang
Gong, Yeyun
Wang, Yanzhi
Shen, Yelong
Machine Learning
Artificial Intelligence
Linear State Space Models (SSMs) offer remarkable performance gains in efficient sequence modeling, with constant inference-time computation and memory complexity. Recent advances, such as Mamba, further enhance SSMs with input-dependent gating and hardware-aware implementations, positioning them as strong alternatives to Transformers for long sequence modeling. However, efficiently scaling the expressive power of SSMs, particularly with Mixture of Experts (MoE), remains challenging, as naive integration attempts often falter or degrade performance. In this work, we introduce Routing Mamba (RoM), a novel approach that scales SSM parameters using sparse mixtures of linear projection experts. By sharing routing decisions between projection layers and lightweight sub-modules within Mamba across experts, RoM leverages synergies among linear projection experts for effective and efficient sparse scaling of Mamba layers. At a scale of 1.3B active parameters (10B total) and 16K training sequence length, RoM achieves language modeling performance equivalent to a dense Mamba model requiring over 2.3x more active parameters, and demonstrates consistent perplexity across context lengths. Experimental results further show RoM effectively scales hybrid language models, yielding a 23% FLOPS saving compared to dense Mamba scaling for similar performance.
title Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.18145