LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Elango, Venmugil, Bhatia, Nidhi, Waleffe, Roger, Shafipour, Rasoul, Asida, Tomer, Khattar, Abhinav, Assaf, Nave, Golub, Maximilian, Guman, Joey, Mitra, Tiyasa, Zhao, Ritchie, Borkar, Ritika, Zilberstein, Ran, Patwary, Mostofa, Shoeybi, Mohammad, Rouhani, Bita
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912848768139264
author Elango, Venmugil
Bhatia, Nidhi
Waleffe, Roger
Shafipour, Rasoul
Asida, Tomer
Khattar, Abhinav
Assaf, Nave
Golub, Maximilian
Guman, Joey
Mitra, Tiyasa
Zhao, Ritchie
Borkar, Ritika
Zilberstein, Ran
Patwary, Mostofa
Shoeybi, Mohammad
Rouhani, Bita
author_facet Elango, Venmugil
Bhatia, Nidhi
Waleffe, Roger
Shafipour, Rasoul
Asida, Tomer
Khattar, Abhinav
Assaf, Nave
Golub, Maximilian
Guman, Joey
Mitra, Tiyasa
Zhao, Ritchie
Borkar, Ritika
Zilberstein, Ran
Patwary, Mostofa
Shoeybi, Mohammad
Rouhani, Bita
contents Mixture of Experts (MoEs) have become a central component of many state-of-the-art open-source and proprietary large language models. Despite their widespread adoption, it remains unclear how close existing MoE architectures are to optimal with respect to inference cost, as measured by accuracy per floating-point operation and per parameter. In this work, we revisit MoE design from a hardware-software co-design perspective, grounded in empirical and theoretical considerations. We characterize key performance bottlenecks across diverse deployment regimes, spanning offline high-throughput execution and online, latency-critical inference. Guided by these insights, we introduce LatentMoE, a new model architecture resulting from systematic design exploration and optimized for maximal accuracy per unit of compute. Empirical design space exploration at scales of up to 95B parameters and over a 1T-token training horizon, together with supporting theoretical analysis, shows that LatentMoE consistently outperforms standard MoE architectures in terms of accuracy per FLOP and per parameter. Given its strong performance, the LatentMoE architecture has been adopted by the flagship Nemotron-3 Super and Ultra models and scaled to substantially larger regimes, including longer token horizons and larger model sizes, as reported in Nvidia et al. (arXiv:2512.20856).
format Preprint
id arxiv_https___arxiv_org_abs_2601_18089
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts
Elango, Venmugil
Bhatia, Nidhi
Waleffe, Roger
Shafipour, Rasoul
Asida, Tomer
Khattar, Abhinav
Assaf, Nave
Golub, Maximilian
Guman, Joey
Mitra, Tiyasa
Zhao, Ritchie
Borkar, Ritika
Zilberstein, Ran
Patwary, Mostofa
Shoeybi, Mohammad
Rouhani, Bita
Machine Learning
Artificial Intelligence
Mixture of Experts (MoEs) have become a central component of many state-of-the-art open-source and proprietary large language models. Despite their widespread adoption, it remains unclear how close existing MoE architectures are to optimal with respect to inference cost, as measured by accuracy per floating-point operation and per parameter. In this work, we revisit MoE design from a hardware-software co-design perspective, grounded in empirical and theoretical considerations. We characterize key performance bottlenecks across diverse deployment regimes, spanning offline high-throughput execution and online, latency-critical inference. Guided by these insights, we introduce LatentMoE, a new model architecture resulting from systematic design exploration and optimized for maximal accuracy per unit of compute. Empirical design space exploration at scales of up to 95B parameters and over a 1T-token training horizon, together with supporting theoretical analysis, shows that LatentMoE consistently outperforms standard MoE architectures in terms of accuracy per FLOP and per parameter. Given its strong performance, the LatentMoE architecture has been adopted by the flagship Nemotron-3 Super and Ultra models and scaled to substantially larger regimes, including longer token horizons and larger model sizes, as reported in Nvidia et al. (arXiv:2512.20856).
title LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2601.18089