GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Yu, Pan, Lehan, Peng, Jie, Tao, Ziyang, Zhu, Hanqi, Zhang, Wuyang, Zhang, Yanyong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913092606099456
author Han, Yu
Pan, Lehan
Peng, Jie
Tao, Ziyang
Zhu, Hanqi
Zhang, Wuyang
Zhang, Yanyong
author_facet Han, Yu
Pan, Lehan
Peng, Jie
Tao, Ziyang
Zhu, Hanqi
Zhang, Wuyang
Zhang, Yanyong
contents Sparse Mixture of Experts (SMoE) enables scalable parameter growth in large language models (LLMs) by selectively activating a subset of experts, and its large parameter count necessitates distributed deployment for inference. However, distributed inference faces a critical dilemma: although communication overhead constitutes the primary bottleneck, reducing it often exacerbates computational load imbalance, leading to resource waste. In this paper, we present GRACE-MoE, which stands for Grouping and Replication with Locality-Aware Routing for SMoE inference. GRACE-MoE is a lossless co-optimization framework that integrates expert grouping to reduce communication and dynamic replication to correct load skew, together with locality-aware routing to resolve replica selection. To underpin this coordinated optimization in multi-node settings, GRACE-MoE adopts a hierarchical sparse communication design that reduces cross-node traffic while implicitly aligning execution across nodes, thereby mitigating synchronization overhead. Experiments on diverse models and multi-node, multi-GPU environments demonstrate that GRACE-MoE efficiently reduces end-to-end inference latency, achieving up to 4.66x speedup over existing systems, and the code will be released upon acceptance.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25041
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
Han, Yu
Pan, Lehan
Peng, Jie
Tao, Ziyang
Zhu, Hanqi
Zhang, Wuyang
Zhang, Yanyong
Distributed, Parallel, and Cluster Computing
Sparse Mixture of Experts (SMoE) enables scalable parameter growth in large language models (LLMs) by selectively activating a subset of experts, and its large parameter count necessitates distributed deployment for inference. However, distributed inference faces a critical dilemma: although communication overhead constitutes the primary bottleneck, reducing it often exacerbates computational load imbalance, leading to resource waste. In this paper, we present GRACE-MoE, which stands for Grouping and Replication with Locality-Aware Routing for SMoE inference. GRACE-MoE is a lossless co-optimization framework that integrates expert grouping to reduce communication and dynamic replication to correct load skew, together with locality-aware routing to resolve replica selection. To underpin this coordinated optimization in multi-node settings, GRACE-MoE adopts a hierarchical sparse communication design that reduces cross-node traffic while implicitly aligning execution across nodes, thereby mitigating synchronization overhead. Experiments on diverse models and multi-node, multi-GPU environments demonstrate that GRACE-MoE efficiently reduces end-to-end inference latency, achieving up to 4.66x speedup over existing systems, and the code will be released upon acceptance.
title GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2509.25041