HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Haoran, Yu, Xianzhi, Zhao, Kang, Bao, Han, Zhan, Zongyuan, Hu, Ting, Liu, Wulong, Yin, Zekun, Li, Xin, Liu, Weiguo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914008580227072
author Lin, Haoran
Yu, Xianzhi
Zhao, Kang
Bao, Han
Zhan, Zongyuan
Hu, Ting
Liu, Wulong
Yin, Zekun
Li, Xin
Liu, Weiguo
author_facet Lin, Haoran
Yu, Xianzhi
Zhao, Kang
Bao, Han
Zhan, Zongyuan
Hu, Ting
Liu, Wulong
Yin, Zekun
Li, Xin
Liu, Weiguo
contents Current inference systems for Mixture-of-Experts (MoE) models primarily employ static parallelization strategies. However, these static approaches cannot consistently achieve optimal performance across different inference scenarios, as they lack the flexibility to adapt to varying computational requirements. In this work, we propose HAP (Hybrid Adaptive Parallelism), a novel method that dynamically selects hybrid parallel strategies to enhance MoE inference efficiency. The fundamental innovation of HAP lies in hierarchically decomposing MoE architectures into two distinct computational modules: the Attention module and the Expert module, each augmented with a specialized inference latency simulation model. This decomposition promotes the construction of a comprehensive search space for seeking model parallel strategies. By leveraging Integer Linear Programming (ILP), HAP could solve the optimal hybrid parallel configurations to maximize inference efficiency under varying computational constraints. Our experiments demonstrate that HAP consistently determines parallel configurations that achieve comparable or superior performance to the TP strategy prevalent in mainstream inference systems. Compared to the TP-based inference, HAP-based inference achieves speedups of 1.68x, 1.77x, and 1.57x on A100, A6000, and V100 GPU platforms, respectively. Furthermore, HAP showcases remarkable generalization capability, maintaining performance effectiveness across diverse MoE model configurations, including Mixtral and Qwen series models.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19373
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
Lin, Haoran
Yu, Xianzhi
Zhao, Kang
Bao, Han
Zhan, Zongyuan
Hu, Ting
Liu, Wulong
Yin, Zekun
Li, Xin
Liu, Weiguo
Distributed, Parallel, and Cluster Computing
Current inference systems for Mixture-of-Experts (MoE) models primarily employ static parallelization strategies. However, these static approaches cannot consistently achieve optimal performance across different inference scenarios, as they lack the flexibility to adapt to varying computational requirements. In this work, we propose HAP (Hybrid Adaptive Parallelism), a novel method that dynamically selects hybrid parallel strategies to enhance MoE inference efficiency. The fundamental innovation of HAP lies in hierarchically decomposing MoE architectures into two distinct computational modules: the Attention module and the Expert module, each augmented with a specialized inference latency simulation model. This decomposition promotes the construction of a comprehensive search space for seeking model parallel strategies. By leveraging Integer Linear Programming (ILP), HAP could solve the optimal hybrid parallel configurations to maximize inference efficiency under varying computational constraints. Our experiments demonstrate that HAP consistently determines parallel configurations that achieve comparable or superior performance to the TP strategy prevalent in mainstream inference systems. Compared to the TP-based inference, HAP-based inference achieves speedups of 1.68x, 1.77x, and 1.57x on A100, A6000, and V100 GPU platforms, respectively. Furthermore, HAP showcases remarkable generalization capability, maintaining performance effectiveness across diverse MoE model configurations, including Mixtral and Qwen series models.
title HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2508.19373