Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wu, Chenpeng, Gu, Qiqi, Shi, Heng, Yao, Jianguo, Guan, Haibing
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913735868678144
author Wu, Chenpeng
Gu, Qiqi
Shi, Heng
Yao, Jianguo
Guan, Haibing
author_facet Wu, Chenpeng
Gu, Qiqi
Shi, Heng
Yao, Jianguo
Guan, Haibing
contents The escalating size of Mixture-of-Experts (MoE) based Large Language Models (LLMs) presents significant computational and memory challenges, necessitating innovative solutions to enhance efficiency without compromising model accuracy. Structured sparsity emerges as a compelling strategy to address these challenges by leveraging the emerging sparse computing hardware. Prior works mainly focus on the sparsity in model parameters, neglecting the inherent sparse patterns in activations. This oversight can lead to additional computational costs associated with activations, potentially resulting in suboptimal performance. This paper presents Samoyeds, an innovative acceleration system for MoE LLMs utilizing Sparse Tensor Cores (SpTCs). Samoyeds is the first to apply sparsity simultaneously to both activations and model parameters. It introduces a bespoke sparse data format tailored for MoE computation and develops a specialized sparse-sparse matrix multiplication kernel. Furthermore, Samoyeds incorporates systematic optimizations specifically designed for the execution of dual-side structured sparse MoE LLMs on SpTCs, further enhancing system performance. Evaluations show that Samoyeds outperforms SOTA works by up to 1.99$\times$ at the kernel level and 1.58$\times$ at the model level. Moreover, it enhances memory efficiency, increasing maximum supported batch sizes by 4.41$\times$ on average. Additionally, Samoyeds surpasses existing SOTA structured sparse solutions in both model accuracy and hardware portability.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10725
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores
Wu, Chenpeng
Gu, Qiqi
Shi, Heng
Yao, Jianguo
Guan, Haibing
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Operating Systems
The escalating size of Mixture-of-Experts (MoE) based Large Language Models (LLMs) presents significant computational and memory challenges, necessitating innovative solutions to enhance efficiency without compromising model accuracy. Structured sparsity emerges as a compelling strategy to address these challenges by leveraging the emerging sparse computing hardware. Prior works mainly focus on the sparsity in model parameters, neglecting the inherent sparse patterns in activations. This oversight can lead to additional computational costs associated with activations, potentially resulting in suboptimal performance. This paper presents Samoyeds, an innovative acceleration system for MoE LLMs utilizing Sparse Tensor Cores (SpTCs). Samoyeds is the first to apply sparsity simultaneously to both activations and model parameters. It introduces a bespoke sparse data format tailored for MoE computation and develops a specialized sparse-sparse matrix multiplication kernel. Furthermore, Samoyeds incorporates systematic optimizations specifically designed for the execution of dual-side structured sparse MoE LLMs on SpTCs, further enhancing system performance. Evaluations show that Samoyeds outperforms SOTA works by up to 1.99$\times$ at the kernel level and 1.58$\times$ at the model level. Moreover, it enhances memory efficiency, increasing maximum supported batch sizes by 4.41$\times$ on average. Additionally, Samoyeds surpasses existing SOTA structured sparse solutions in both model accuracy and hardware portability.
title Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Operating Systems
url https://arxiv.org/abs/2503.10725