MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
Fuente:
arXiv
Saved in:
| Main Authors: | Zibakhsh, Soheil, Samragh, Mohammad, Nishu, Kumari, Hannah, Lauren, Kundu, Arnav, Cho, Minsik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
by: Hannah, Lauren. A, et al.
Published: (2025)
by: Hannah, Lauren. A, et al.
Published: (2025)
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
by: Samragh, Mohammad, et al.
Published: (2025)
by: Samragh, Mohammad, et al.
Published: (2025)
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
by: Kim, Han-Byul, et al.
Published: (2025)
by: Kim, Han-Byul, et al.
Published: (2025)
Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
by: Bhendawade, Nikhil, et al.
Published: (2025)
by: Bhendawade, Nikhil, et al.
Published: (2025)
SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
by: Nishu, Kumari, et al.
Published: (2024)
by: Nishu, Kumari, et al.
Published: (2024)
SpecMD: A Comprehensive Study On Speculative Expert Prefetching
by: Hoang, Duc, et al.
Published: (2026)
by: Hoang, Duc, et al.
Published: (2026)
Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder
by: Bittar, Alexandre, et al.
Published: (2023)
by: Bittar, Alexandre, et al.
Published: (2023)
Gleaning Insights From Research on Evaluation (RoE) PhD Dissertations
by: Jennifer P. Villalobos, et al.
Published: (2025)
by: Jennifer P. Villalobos, et al.
Published: (2025)
KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
by: Cho, Minsik, et al.
Published: (2024)
by: Cho, Minsik, et al.
Published: (2024)
MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers
by: Jaiswal, Ajay, et al.
Published: (2026)
by: Jaiswal, Ajay, et al.
Published: (2026)
R2 Loss: Range Restriction Loss for Model Compression and Quantization
by: Kundu, Arnav, et al.
Published: (2023)
by: Kundu, Arnav, et al.
Published: (2023)
From Dense to Dynamic: Token-Difficulty Driven MoEfication of Pre-Trained LLMs
by: Nishu, Kumari, et al.
Published: (2025)
by: Nishu, Kumari, et al.
Published: (2025)
Off-diagonally symmetric alternating sign matrices
by: Kumari, Nishu
Published: (2025)
by: Kumari, Nishu
Published: (2025)
A determinantal formula for orthosymplectic Schur functions
by: Kumari, Nishu
Published: (2024)
by: Kumari, Nishu
Published: (2024)
Skew hook Schur functions and the cyclic sieving phenomenon
by: Kumari, Nishu
Published: (2022)
by: Kumari, Nishu
Published: (2022)
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments
by: Kim, Minsoo, et al.
Published: (2025)
by: Kim, Minsoo, et al.
Published: (2025)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
by: Samragh, Mohammad, et al.
Published: (2024)
by: Samragh, Mohammad, et al.
Published: (2024)
LayerCollapse: Adaptive compression of neural networks
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2023)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2023)
Towards Low-bit Communication for Tensor Parallel LLM Inference
by: Dong, Harry, et al.
Published: (2024)
by: Dong, Harry, et al.
Published: (2024)
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core
by: Liu, Dennis, et al.
Published: (2025)
by: Liu, Dennis, et al.
Published: (2025)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
by: Luo, Jiajun, et al.
Published: (2024)
by: Luo, Jiajun, et al.
Published: (2024)
The Weak Form Is Stronger Than You Think
by: Messenger, Daniel A., et al.
Published: (2024)
by: Messenger, Daniel A., et al.
Published: (2024)
Beyond Perplexity: A Lightweight Benchmark for Knowledge Retention in Supervised Fine-Tuning
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2026)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2026)
MergeGuard: Efficient Thwarting of Trojan Attacks in Machine Learning Models
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism
by: Cui, Chenwei, et al.
Published: (2026)
by: Cui, Chenwei, et al.
Published: (2026)
DOT-MoE: Differentiable Optimal Transport for MoEfication
by: Bamba, Udbhav, et al.
Published: (2026)
by: Bamba, Udbhav, et al.
Published: (2026)
Learning How Much to Think: Difficulty-Aware Dynamic MoEs for Graph Node Classification
by: Zhou, Jiajun, et al.
Published: (2026)
by: Zhou, Jiajun, et al.
Published: (2026)
Limit profile for the transpose top-2 with random shuffle
by: Ghosh, Subhajit, et al.
Published: (2024)
by: Ghosh, Subhajit, et al.
Published: (2024)
Murnaghan--Nakayama rules for symplectic, orthogonal and orthosymplectic Schur functions
by: Kumari, Nishu, et al.
Published: (2024)
by: Kumari, Nishu, et al.
Published: (2024)
Further results for classical and universal characters twisted by roots of unity
by: Ayyer, Arvind, et al.
Published: (2024)
by: Ayyer, Arvind, et al.
Published: (2024)
RoE-FND: A Case-Based Reasoning Approach with Dual Verification for Fake News Detection via LLMs
by: Yang, Yuzhou, et al.
Published: (2025)
by: Yang, Yuzhou, et al.
Published: (2025)
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference
by: Sun, Xun, et al.
Published: (2026)
by: Sun, Xun, et al.
Published: (2026)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
by: Pan, Xinglin, et al.
Published: (2025)
by: Pan, Xinglin, et al.
Published: (2025)
Sparse Crosscoders for diffing MoEs and Dense models
by: Chaudhari, Marmik, et al.
Published: (2026)
by: Chaudhari, Marmik, et al.
Published: (2026)
LiveTune: Dynamic Parameter Tuning for Feedback-Driven Optimization
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2023)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2023)
ForTIFAI: Fending Off Recursive Training Induced Failure for AI Model Collapse
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
Stronger Than You Think: Benchmarking Weak Supervision on Realistic Tasks
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
by: Cao, Shiyi, et al.
Published: (2024)
by: Cao, Shiyi, et al.
Published: (2024)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
by: Qian, Yulei, et al.
Published: (2024)
by: Qian, Yulei, et al.
Published: (2024)
Similar Items
-
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
by: Hannah, Lauren. A, et al.
Published: (2025) -
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
by: Samragh, Mohammad, et al.
Published: (2025) -
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
by: Kim, Han-Byul, et al.
Published: (2025) -
Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
by: Bhendawade, Nikhil, et al.
Published: (2025) -
SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
by: Nishu, Kumari, et al.
Published: (2024)