Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Haoyuan, Chen, Haoxing, Chen, Xiaodong, Zhou, Zhanchao, Chen, Tieyuan, Zhuang, Yihong, Lu, Guoshan, Huang, Zenan, Zhao, Junbo, Liu, Lin, Lan, Zhenzhong, Yu, Bei, Li, Jianguo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913984295206912
author Wu, Haoyuan
Chen, Haoxing
Chen, Xiaodong
Zhou, Zhanchao
Chen, Tieyuan
Zhuang, Yihong
Lu, Guoshan
Huang, Zenan
Zhao, Junbo
Liu, Lin
Lan, Zhenzhong
Yu, Bei
Li, Jianguo
author_facet Wu, Haoyuan
Chen, Haoxing
Chen, Xiaodong
Zhou, Zhanchao
Chen, Tieyuan
Zhuang, Yihong
Lu, Guoshan
Huang, Zenan
Zhao, Junbo
Liu, Lin
Lan, Zhenzhong
Yu, Bei
Li, Jianguo
contents The Mixture of Experts (MoE) architecture is a cornerstone of modern state-of-the-art (SOTA) large language models (LLMs). MoE models facilitate scalability by enabling sparse parameter activation. However, traditional MoE architecture uses homogeneous experts of a uniform size, activating a fixed number of parameters irrespective of input complexity and thus limiting computational efficiency. To overcome this limitation, we introduce Grove MoE, a novel architecture incorporating experts of varying sizes, inspired by the heterogeneous big.LITTLE CPU architecture. This architecture features novel adjugate experts with a dynamic activation mechanism, enabling model capacity expansion while maintaining manageable computational overhead. Building on this architecture, we present GroveMoE-Base and GroveMoE-Inst, 33B-parameter LLMs developed by applying an upcycling strategy to the Qwen3-30B-A3B-Base model during mid-training and post-training. GroveMoE models dynamically activate 3.14-3.28B parameters based on token complexity and achieve performance comparable to SOTA open-source models of similar or even larger size.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07785
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
Wu, Haoyuan
Chen, Haoxing
Chen, Xiaodong
Zhou, Zhanchao
Chen, Tieyuan
Zhuang, Yihong
Lu, Guoshan
Huang, Zenan
Zhao, Junbo
Liu, Lin
Lan, Zhenzhong
Yu, Bei
Li, Jianguo
Computation and Language
The Mixture of Experts (MoE) architecture is a cornerstone of modern state-of-the-art (SOTA) large language models (LLMs). MoE models facilitate scalability by enabling sparse parameter activation. However, traditional MoE architecture uses homogeneous experts of a uniform size, activating a fixed number of parameters irrespective of input complexity and thus limiting computational efficiency. To overcome this limitation, we introduce Grove MoE, a novel architecture incorporating experts of varying sizes, inspired by the heterogeneous big.LITTLE CPU architecture. This architecture features novel adjugate experts with a dynamic activation mechanism, enabling model capacity expansion while maintaining manageable computational overhead. Building on this architecture, we present GroveMoE-Base and GroveMoE-Inst, 33B-parameter LLMs developed by applying an upcycling strategy to the Qwen3-30B-A3B-Base model during mid-training and post-training. GroveMoE models dynamically activate 3.14-3.28B parameters based on token complexity and achieve performance comparable to SOTA open-source models of similar or even larger size.
title Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
topic Computation and Language
url https://arxiv.org/abs/2508.07785