LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Elango, Venmugil, Bhatia, Nidhi, Waleffe, Roger, Shafipour, Rasoul, Asida, Tomer, Khattar, Abhinav, Assaf, Nave, Golub, Maximilian, Guman, Joey, Mitra, Tiyasa, Zhao, Ritchie, Borkar, Ritika, Zilberstein, Ran, Patwary, Mostofa, Shoeybi, Mohammad, Rouhani, Bita |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
by: Iso, Hayate, et al.
Published: (2026)
by: Iso, Hayate, et al.
Published: (2026)
Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
by: Bhatia, Nidhi, et al.
Published: (2025)
by: Bhatia, Nidhi, et al.
Published: (2025)
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
by: Mitra, Tiyasa, et al.
Published: (2025)
by: Mitra, Tiyasa, et al.
Published: (2025)
ATTENTION2D: Communication Efficient Distributed Self-Attention Mechanism
by: Elango, Venmugil
Published: (2025)
by: Elango, Venmugil
Published: (2025)
PaSE: Parallelization Strategies for Efficient DNN Training
by: Elango, Venmugil
Published: (2024)
by: Elango, Venmugil
Published: (2024)
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
by: Abramovich, Talor, et al.
Published: (2026)
by: Abramovich, Talor, et al.
Published: (2026)
Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
by: Parmar, Jupinder, et al.
Published: (2024)
by: Parmar, Jupinder, et al.
Published: (2024)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
by: Feng, Steven, et al.
Published: (2024)
by: Feng, Steven, et al.
Published: (2024)
Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
by: Akter, Syeda Nahida, et al.
Published: (2025)
by: Akter, Syeda Nahida, et al.
Published: (2025)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
by: Yu, Yanpeng, et al.
Published: (2025)
by: Yu, Yanpeng, et al.
Published: (2025)
FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs
by: Akter, Syeda Nahida, et al.
Published: (2024)
by: Akter, Syeda Nahida, et al.
Published: (2024)
RLP: Reinforcement as a Pretraining Objective
by: Hatamizadeh, Ali, et al.
Published: (2025)
by: Hatamizadeh, Ali, et al.
Published: (2025)
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
by: Su, Dan, et al.
Published: (2024)
by: Su, Dan, et al.
Published: (2024)
Data, Data Everywhere: A Guide for Pretraining Dataset Construction
by: Parmar, Jupinder, et al.
Published: (2024)
by: Parmar, Jupinder, et al.
Published: (2024)
Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques
by: Javidnia, Neusha, et al.
Published: (2025)
by: Javidnia, Neusha, et al.
Published: (2025)
Compact Language Models via Pruning and Knowledge Distillation
by: Muralidharan, Saurav, et al.
Published: (2024)
by: Muralidharan, Saurav, et al.
Published: (2024)
Emission Distribution for the quantas of Maxwell-Chern-Simon Gauge Field coupled to External Current
by: Kar, Tiyasa
Published: (2021)
by: Kar, Tiyasa
Published: (2021)
Extending Puzzle for Mixture-of-Experts Reasoning Models with Application to GPT-OSS Acceleration
by: Bercovich, Akhiad, et al.
Published: (2026)
by: Bercovich, Akhiad, et al.
Published: (2026)
Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
by: Lu, Ximing, et al.
Published: (2025)
by: Lu, Ximing, et al.
Published: (2025)
Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
by: Akter, Syeda Nahida, et al.
Published: (2025)
by: Akter, Syeda Nahida, et al.
Published: (2025)
Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism
by: Cui, Chenwei, et al.
Published: (2026)
by: Cui, Chenwei, et al.
Published: (2026)
FALCON: FLOP-Aware Combinatorial Optimization for Neural Network Pruning
by: Meng, Xiang, et al.
Published: (2024)
by: Meng, Xiang, et al.
Published: (2024)
Anonymous Survey on On-demand e-exams WS 2024-25
by: Patwary, Mozaher
Published: (2026)
by: Patwary, Mozaher
Published: (2026)
An Investigation into Emotional Intelligence, Foreign Language Anxiety and Empathy through a Cognitive-Affective Course in an EFL Context
by: Ali Rouhani
Published: (2008)
by: Ali Rouhani
Published: (2008)
ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration
by: Ai, Mengting, et al.
Published: (2025)
by: Ai, Mengting, et al.
Published: (2025)
Sparse-IFT: Sparse Iso-FLOP Transformations for Maximizing Training Efficiency
by: Thangarasa, Vithursan, et al.
Published: (2023)
by: Thangarasa, Vithursan, et al.
Published: (2023)
Outcome Logic: A Unified Approach to the Metatheory of Program Logics with Branching Effects
by: Zilberstein, Noam
Published: (2024)
by: Zilberstein, Noam
Published: (2024)
SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random Generators
by: Shafipour, Rasoul, et al.
Published: (2024)
by: Shafipour, Rasoul, et al.
Published: (2024)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)
by: Cheng, Long, et al.
Published: (2026)
FLOP-Efficient Training: Early Stopping Based on Test-Time Compute Awareness
by: Amer, Hossam, et al.
Published: (2026)
by: Amer, Hossam, et al.
Published: (2026)
Problems with Chinchilla Approach 2: Systematic Biases in IsoFLOP Parabola Fits
by: Czech, Eric, et al.
Published: (2026)
by: Czech, Eric, et al.
Published: (2026)
Mixed Sparsity Training: Achieving 4$\times$ FLOP Reduction for Transformer Pretraining
by: Hu, Pihe, et al.
Published: (2024)
by: Hu, Pihe, et al.
Published: (2024)
Neural networks can be FLOP-efficient integrators of 1D oscillatory integrands
by: Sinha, Anshuman, et al.
Published: (2024)
by: Sinha, Anshuman, et al.
Published: (2024)
Thermodynamic Characteristics of a Fermi Gas with an Invariant Energy Scale and its Astrophysical Implications
by: Kar, Tiyasa, et al.
Published: (2026)
by: Kar, Tiyasa, et al.
Published: (2026)
Stochastic Approximation with Two Time Scales: The General Case
by: Borkar, Vivek S
Published: (2024)
by: Borkar, Vivek S
Published: (2024)
Multi-Agent Evolve: LLM Self-Improve through Co-evolution
by: Chen, Yixing, et al.
Published: (2025)
by: Chen, Yixing, et al.
Published: (2025)
Quantitative (Chlorophyll-a) and qualitative (species composition) seasonal fluctuations of phytoplankton in Lavan coastal waters (North of the Persian Gulf)
by: Rouhani Ghadikolaei, K.
Published: (2001)
by: Rouhani Ghadikolaei, K.
Published: (2001)
LAS ESTRATEGIAS DE APRENDIZAJE DESDE UNA DIDÁCTICA DESARROLLADORA
by: José Zilberstein Toruncha
Published: (2014)
by: José Zilberstein Toruncha
Published: (2014)
Similar Items
-
Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
by: Iso, Hayate, et al.
Published: (2026) -
Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
by: Bhatia, Nidhi, et al.
Published: (2025) -
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
by: Mitra, Tiyasa, et al.
Published: (2025) -
ATTENTION2D: Communication Efficient Distributed Self-Attention Mechanism
by: Elango, Venmugil
Published: (2025) -
PaSE: Parallelization Strategies for Efficient DNN Training
by: Elango, Venmugil
Published: (2024)