Janus: Disaggregating Attention and Experts for Scalable MoE Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhexiang, Wang, Ye, Zhao, Yumiao, Xiao, Jiayu, Yang, Qianjing, Wang, Xiangyu, Jiang, Jingzhe, Weng, Qizhen, Chen, Ruichuan, Shi, Shaohuai, Toosi, Adel N., Chen, Yin, Yu, Minchen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917441545699328
author Zhang, Zhexiang
Wang, Ye
Zhao, Yumiao
Xiao, Jiayu
Yang, Qianjing
Wang, Xiangyu
Jiang, Jingzhe
Weng, Qizhen
Chen, Ruichuan
Shi, Shaohuai
Toosi, Adel N.
Chen, Yin
Yu, Minchen
author_facet Zhang, Zhexiang
Wang, Ye
Zhao, Yumiao
Xiao, Jiayu
Yang, Qianjing
Wang, Xiangyu
Jiang, Jingzhe
Weng, Qizhen
Chen, Ruichuan
Shi, Shaohuai
Toosi, Adel N.
Chen, Yin
Yu, Minchen
contents Serving large Mixture-of-Experts (MoE) models is challenging because of their large memory footprints, heterogeneous resource demands, and highly dynamic inference workloads. Most existing MoE inference systems deploy the entire model as a monolithic unit, forcing attention and MoE layers to share the same resource configuration despite their different scaling behaviors and resource bottlenecks. Such coarse-grained provisioning leads to resource inefficiency and suboptimal performance. We present JANUS, a scalable and resource-efficient MoE inference system built around three key principles. First, JANUS disaggregates attention and MoE layers onto separate GPU worker pools, enabling independent resource provisioning for the two layer types, and uses an adaptive two-phase communication mechanism for low-latency data exchange. Second, because MoE-layer execution is often memory-bound and highly sensitive to activated-expert imbalance, JANUS introduces a lightweight, microsecond-scale activation scheduler that balances per-layer activated experts across MoE instances to reduce inference latency. Third, JANUS employs a fine-grained, SLO-aware resource scaling scheme that jointly selects attention resources, MoE resources, and expert placement to minimize GPU cost under token-level SLOs. Evaluation shows that JANUS improves per-GPU throughput by up to 4.7x over state-of-the-art MoE inference baselines while satisfying token-level latency SLOs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13525
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Janus: Disaggregating Attention and Experts for Scalable MoE Inference
Zhang, Zhexiang
Wang, Ye
Zhao, Yumiao
Xiao, Jiayu
Yang, Qianjing
Wang, Xiangyu
Jiang, Jingzhe
Weng, Qizhen
Chen, Ruichuan
Shi, Shaohuai
Toosi, Adel N.
Chen, Yin
Yu, Minchen
Distributed, Parallel, and Cluster Computing
Serving large Mixture-of-Experts (MoE) models is challenging because of their large memory footprints, heterogeneous resource demands, and highly dynamic inference workloads. Most existing MoE inference systems deploy the entire model as a monolithic unit, forcing attention and MoE layers to share the same resource configuration despite their different scaling behaviors and resource bottlenecks. Such coarse-grained provisioning leads to resource inefficiency and suboptimal performance. We present JANUS, a scalable and resource-efficient MoE inference system built around three key principles. First, JANUS disaggregates attention and MoE layers onto separate GPU worker pools, enabling independent resource provisioning for the two layer types, and uses an adaptive two-phase communication mechanism for low-latency data exchange. Second, because MoE-layer execution is often memory-bound and highly sensitive to activated-expert imbalance, JANUS introduces a lightweight, microsecond-scale activation scheduler that balances per-layer activated experts across MoE instances to reduce inference latency. Third, JANUS employs a fine-grained, SLO-aware resource scaling scheme that jointly selects attention resources, MoE resources, and expert placement to minimize GPU cost under token-level SLOs. Evaluation shows that JANUS improves per-GPU throughput by up to 4.7x over state-of-the-art MoE inference baselines while satisfying token-level latency SLOs.
title Janus: Disaggregating Attention and Experts for Scalable MoE Inference
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2512.13525