GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Chen, Cai, Yuhang, Liu, Deyi, Ma, Jin, Ma, Yiyuan, Yang, Yuan, Liu, Jing, Zeng, Yutao, Zhou, Xun, Qiao, Siyuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914093831553024
author Zheng, Chen
Cai, Yuhang
Liu, Deyi
Ma, Jin
Ma, Yiyuan
Yang, Yuan
Liu, Jing
Zeng, Yutao
Zhou, Xun
Qiao, Siyuan
author_facet Zheng, Chen
Cai, Yuhang
Liu, Deyi
Ma, Jin
Ma, Yiyuan
Yang, Yuan
Liu, Jing
Zeng, Yutao
Zhou, Xun
Qiao, Siyuan
contents Modern large language models leverage Mixture-of-Experts (MoE) architectures for efficient scaling, but face a critical challenge: functionally similar experts are often selected simultaneously, creating redundant computation and limiting effective model capacity. Existing auxiliary balance loss methods improve token distribution but fail to address the underlying expert diversity problem. We introduce GatePro, a novel parameter-free method that directly promotes expert selection diversity. GatePro identifies the most similar expert pairs and introduces localized competition mechanisms, preventing redundant expert co-activation while maintaining natural expert specialization. Our comprehensive evaluation demonstrates GatePro's effectiveness across model scales and benchmarks. Analysis demonstrates GatePro's ability to achieve enhanced expert diversity, where experts develop more distinct and complementary capabilities, avoiding functional redundancy. This approach can be deployed hot-swappable during any training phase without additional learnable parameters, offering a practical solution for improving MoE effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13079
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
Zheng, Chen
Cai, Yuhang
Liu, Deyi
Ma, Jin
Ma, Yiyuan
Yang, Yuan
Liu, Jing
Zeng, Yutao
Zhou, Xun
Qiao, Siyuan
Computation and Language
Machine Learning
Modern large language models leverage Mixture-of-Experts (MoE) architectures for efficient scaling, but face a critical challenge: functionally similar experts are often selected simultaneously, creating redundant computation and limiting effective model capacity. Existing auxiliary balance loss methods improve token distribution but fail to address the underlying expert diversity problem. We introduce GatePro, a novel parameter-free method that directly promotes expert selection diversity. GatePro identifies the most similar expert pairs and introduces localized competition mechanisms, preventing redundant expert co-activation while maintaining natural expert specialization. Our comprehensive evaluation demonstrates GatePro's effectiveness across model scales and benchmarks. Analysis demonstrates GatePro's ability to achieve enhanced expert diversity, where experts develop more distinct and complementary capabilities, avoiding functional redundancy. This approach can be deployed hot-swappable during any training phase without additional learnable parameters, offering a practical solution for improving MoE effectiveness.
title GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.13079