MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gavhane, Nishant, Mehrotra, Arush, Chawla, Rohit, Proenca, Peter
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914001862000640
author Gavhane, Nishant
Mehrotra, Arush
Chawla, Rohit
Proenca, Peter
author_facet Gavhane, Nishant
Mehrotra, Arush
Chawla, Rohit
Proenca, Peter
contents The deployment of large-scale Mixture-of-Experts (MoE) models on edge devices presents significant challenges due to memory constraints. While MoE architectures enable efficient utilization of computational resources by activating only a subset of experts per inference, they require careful memory management to operate efficiently in resource-constrained environments. Traditional heuristic-based expert caching strategies such as MoE-Infinity struggle to maintain high cache hit rates as models parameters scale. In this work, we introduce MoE-Beyond, a learning-based expert activation predictor trained to predict expert activations during autoregressive decoding. By framing the task as a multi-label sequence prediction problem, we train a lightweight transformer model on 66 million expert activation traces extracted from LDJnr-Puffin dataset [5] using DeepSeek-V2-Chat-Lite MoE. Our predictor generalizes effectively across unseen prompts from WebGLM-QA dataset [6], achieving 97.5% accuracy and an 86.6% F1-score. Simulation results show that MoE-Beyond improves GPU cache hit rate from 17% to 72% when only 10% of experts fit in GPU cache, outperforming heuristic baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17137
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices
Gavhane, Nishant
Mehrotra, Arush
Chawla, Rohit
Proenca, Peter
Machine Learning
The deployment of large-scale Mixture-of-Experts (MoE) models on edge devices presents significant challenges due to memory constraints. While MoE architectures enable efficient utilization of computational resources by activating only a subset of experts per inference, they require careful memory management to operate efficiently in resource-constrained environments. Traditional heuristic-based expert caching strategies such as MoE-Infinity struggle to maintain high cache hit rates as models parameters scale. In this work, we introduce MoE-Beyond, a learning-based expert activation predictor trained to predict expert activations during autoregressive decoding. By framing the task as a multi-label sequence prediction problem, we train a lightweight transformer model on 66 million expert activation traces extracted from LDJnr-Puffin dataset [5] using DeepSeek-V2-Chat-Lite MoE. Our predictor generalizes effectively across unseen prompts from WebGLM-QA dataset [6], achieving 97.5% accuracy and an 86.6% F1-score. Simulation results show that MoE-Beyond improves GPU cache hit rate from 17% to 72% when only 10% of experts fit in GPU cache, outperforming heuristic baselines.
title MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices
topic Machine Learning
url https://arxiv.org/abs/2508.17137