Optimizing MoE Routers: Design, Implementation, and Evaluation in Transformer Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Harvey, Daniel Fidel, Weale, George, Yilmaz, Berk
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908414634885120
author Harvey, Daniel Fidel
Weale, George
Yilmaz, Berk
author_facet Harvey, Daniel Fidel
Weale, George
Yilmaz, Berk
contents Mixture of Experts (MoE) architectures increase large language model scalability, yet their performance depends on the router module that moves tokens to specialized experts. Bad routing can load imbalance and reduced accuracy. This project designed and implemented different router architectures within Transformer models to fix these limitations. We experimented with six distinct router variants Linear, Attention, Multi-Layer Perceptron (MLP), Hybrid, Hash, and our new MLP-Hadamard. We characterized these routers using BERT and the Qwen1.5-MoE model, looking at parameter efficiency, inference latency, routing entropy, and expert utilization patterns. Our evaluations showed distinct trade-offs: Linear routers offer speed, while MLP and Attention routers provide greater expressiveness. The MLP-Hadamard router shows a unique capability for structured, sparse routing. We successfully replaced and fine-tuned custom routers within the complex, quantized Qwen1.5-MoE model. This work provides a comparative analysis of MoE router designs and offers insights into optimizing their performance for efficient and effective large-scale model deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16419
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimizing MoE Routers: Design, Implementation, and Evaluation in Transformer Models
Harvey, Daniel Fidel
Weale, George
Yilmaz, Berk
Machine Learning
Artificial Intelligence
68T07, 68T45
Mixture of Experts (MoE) architectures increase large language model scalability, yet their performance depends on the router module that moves tokens to specialized experts. Bad routing can load imbalance and reduced accuracy. This project designed and implemented different router architectures within Transformer models to fix these limitations. We experimented with six distinct router variants Linear, Attention, Multi-Layer Perceptron (MLP), Hybrid, Hash, and our new MLP-Hadamard. We characterized these routers using BERT and the Qwen1.5-MoE model, looking at parameter efficiency, inference latency, routing entropy, and expert utilization patterns. Our evaluations showed distinct trade-offs: Linear routers offer speed, while MLP and Attention routers provide greater expressiveness. The MLP-Hadamard router shows a unique capability for structured, sparse routing. We successfully replaced and fine-tuned custom routers within the complex, quantized Qwen1.5-MoE model. This work provides a comparative analysis of MoE router designs and offers insights into optimizing their performance for efficient and effective large-scale model deployment.
title Optimizing MoE Routers: Design, Implementation, and Evaluation in Transformer Models
topic Machine Learning
Artificial Intelligence
68T07, 68T45
url https://arxiv.org/abs/2506.16419