Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nakamura, Taishi, Ishikawa, Satoki, Kawamura, Masaki, Okamoto, Takumi, Nohara, Daisuke, Suzuki, Jun, Yokota, Rio
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917301057486848
author Nakamura, Taishi
Ishikawa, Satoki
Kawamura, Masaki
Okamoto, Takumi
Nohara, Daisuke
Suzuki, Jun
Yokota, Rio
author_facet Nakamura, Taishi
Ishikawa, Satoki
Kawamura, Masaki
Okamoto, Takumi
Nohara, Daisuke
Suzuki, Jun
Yokota, Rio
contents Empirical scaling laws have driven the evolution of large language models (LLMs), yet their coefficients shift whenever the model architecture or data pipeline changes. Mixture-of-Experts (MoE) models, now standard in state-of-the-art systems, introduce a new sparsity dimension that current dense-model frontiers overlook. We investigate how MoE sparsity influences two distinct capability regimes: memorization skills and reasoning skills. By training MoE families that vary total parameters, active parameters, and top-$k$ routing under fixed compute budgets, we disentangle pre-training loss from downstream accuracy. Our results reveal two principles. First, Active FLOPs: models with identical training loss but greater active compute achieve higher reasoning accuracy. Second, Total tokens per parameter (TPP): memorization tasks improve with more parameters, while reasoning tasks benefit from optimal TPP, indicating that reasoning is data-hungry. Neither reinforcement learning post-training (GRPO) nor increased test-time compute alters these trends. We therefore argue that optimal MoE sparsity must be determined jointly by active FLOPs and TPP, revising the classical picture of compute-optimal scaling. Our model checkpoints, code and logs are open-source at https://github.com/rioyokotalab/optimal-sparsity.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18672
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
Nakamura, Taishi
Ishikawa, Satoki
Kawamura, Masaki
Okamoto, Takumi
Nohara, Daisuke
Suzuki, Jun
Yokota, Rio
Machine Learning
Artificial Intelligence
Computation and Language
Empirical scaling laws have driven the evolution of large language models (LLMs), yet their coefficients shift whenever the model architecture or data pipeline changes. Mixture-of-Experts (MoE) models, now standard in state-of-the-art systems, introduce a new sparsity dimension that current dense-model frontiers overlook. We investigate how MoE sparsity influences two distinct capability regimes: memorization skills and reasoning skills. By training MoE families that vary total parameters, active parameters, and top-$k$ routing under fixed compute budgets, we disentangle pre-training loss from downstream accuracy. Our results reveal two principles. First, Active FLOPs: models with identical training loss but greater active compute achieve higher reasoning accuracy. Second, Total tokens per parameter (TPP): memorization tasks improve with more parameters, while reasoning tasks benefit from optimal TPP, indicating that reasoning is data-hungry. Neither reinforcement learning post-training (GRPO) nor increased test-time compute alters these trends. We therefore argue that optimal MoE sparsity must be determined jointly by active FLOPs and TPP, revising the classical picture of compute-optimal scaling. Our model checkpoints, code and logs are open-source at https://github.com/rioyokotalab/optimal-sparsity.
title Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.18672