MH-MoE: Multi-Head Mixture-of-Experts

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Huang, Shaohan, Wu, Xun, Ma, Shuming, Wei, Furu
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910720361234432
author Huang, Shaohan
Wu, Xun
Ma, Shuming
Wei, Furu
author_facet Huang, Shaohan
Wu, Xun
Ma, Shuming
Wei, Furu
contents Multi-Head Mixture-of-Experts (MH-MoE) demonstrates superior performance by using the multi-head mechanism to collectively attend to information from various representation spaces within different experts. In this paper, we present a novel implementation of MH-MoE that maintains both FLOPs and parameter parity with sparse Mixture of Experts models. Experimental results on language models show that the new implementation yields quality improvements over both vanilla MoE and fine-grained MoE models. Additionally, our experiments demonstrate that MH-MoE is compatible with 1-bit Large Language Models (LLMs) such as BitNet.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16205
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MH-MoE: Multi-Head Mixture-of-Experts
Huang, Shaohan
Wu, Xun
Ma, Shuming
Wei, Furu
Computation and Language
Multi-Head Mixture-of-Experts (MH-MoE) demonstrates superior performance by using the multi-head mechanism to collectively attend to information from various representation spaces within different experts. In this paper, we present a novel implementation of MH-MoE that maintains both FLOPs and parameter parity with sparse Mixture of Experts models. Experimental results on language models show that the new implementation yields quality improvements over both vanilla MoE and fine-grained MoE models. Additionally, our experiments demonstrate that MH-MoE is compatible with 1-bit Large Language Models (LLMs) such as BitNet.
title MH-MoE: Multi-Head Mixture-of-Experts
topic Computation and Language
url https://arxiv.org/abs/2411.16205