MoE-PHDS: One MoE checkpoint for flexible runtime sparsity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hannah, Lauren. A, Zibakhsh, Soheil, Nishu, Kumari, Kundu, Arnav, Razlighi, Mohammad Samragh, Farajtabar, Mehrdad, Cho, Minsik
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914059079647232
author Hannah, Lauren. A
Zibakhsh, Soheil
Nishu, Kumari
Kundu, Arnav
Razlighi, Mohammad Samragh
Farajtabar, Mehrdad
Cho, Minsik
author_facet Hannah, Lauren. A
Zibakhsh, Soheil
Nishu, Kumari
Kundu, Arnav
Razlighi, Mohammad Samragh
Farajtabar, Mehrdad
Cho, Minsik
contents Sparse Mixtures of Experts (MoEs) are typically trained to operate at a fixed sparsity level, e.g. $k$ in a top-$k$ gating function. This global sparsity level determines an operating point on the accuracy/latency curve; currently, meeting multiple efficiency targets means training and maintaining multiple models. This practice complicates serving, increases training and maintenance costs, and limits flexibility in meeting diverse latency, efficiency, and energy requirements. We show that pretrained MoEs are more robust to runtime sparsity shifts than commonly assumed, and introduce MoE-PHDS ({\bf P}ost {\bf H}oc {\bf D}eclared {\bf S}parsity), a lightweight SFT method that turns a single checkpoint into a global sparsity control surface. PHDS mixes training across sparsity levels and anchors with a short curriculum at high sparsity, requiring no architectural changes. The result is predictable accuracy/latency tradeoffs from one model: practitioners can ``dial $k$'' at inference time without swapping checkpoints, changing architecture, or relying on token-level heuristics. Experiments on OLMoE-1B-7B-0125, Qwen1.5-MoE-A2.7B, and proprietary models fit on multiple operating points show that PHDS matches or exceeds well-specified oracle models, improves cross-sparsity agreement by up to 22\% vs. well-specified oracle models, and enables simplified, flexible runtime MoE deployment by making global sparsity a first-class serving primitive.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23012
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
Hannah, Lauren. A
Zibakhsh, Soheil
Nishu, Kumari
Kundu, Arnav
Razlighi, Mohammad Samragh
Farajtabar, Mehrdad
Cho, Minsik
Machine Learning
Artificial Intelligence
Sparse Mixtures of Experts (MoEs) are typically trained to operate at a fixed sparsity level, e.g. $k$ in a top-$k$ gating function. This global sparsity level determines an operating point on the accuracy/latency curve; currently, meeting multiple efficiency targets means training and maintaining multiple models. This practice complicates serving, increases training and maintenance costs, and limits flexibility in meeting diverse latency, efficiency, and energy requirements. We show that pretrained MoEs are more robust to runtime sparsity shifts than commonly assumed, and introduce MoE-PHDS ({\bf P}ost {\bf H}oc {\bf D}eclared {\bf S}parsity), a lightweight SFT method that turns a single checkpoint into a global sparsity control surface. PHDS mixes training across sparsity levels and anchors with a short curriculum at high sparsity, requiring no architectural changes. The result is predictable accuracy/latency tradeoffs from one model: practitioners can ``dial $k$'' at inference time without swapping checkpoints, changing architecture, or relying on token-level heuristics. Experiments on OLMoE-1B-7B-0125, Qwen1.5-MoE-A2.7B, and proprietary models fit on multiple operating points show that PHDS matches or exceeds well-specified oracle models, improves cross-sparsity agreement by up to 22\% vs. well-specified oracle models, and enables simplified, flexible runtime MoE deployment by making global sparsity a first-class serving primitive.
title MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.23012