MoE-Spec: Expert Budgeting for Efficient Speculative Decoding
Fuente:
arXiv
Guardado en:
| Autores principales: | McDanel, Bradley, Li, Steven, Surineni, Sruthikesh, Khaitan, Harshit |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
por: McDanel, Bradley, et al.
Publicado: (2026)
por: McDanel, Bradley, et al.
Publicado: (2026)
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
por: McDanel, Bradley
Publicado: (2024)
por: McDanel, Bradley
Publicado: (2024)
DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
por: Yu, Fengze, et al.
Publicado: (2025)
por: Yu, Fengze, et al.
Publicado: (2025)
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding
por: Bang, Jehyeon, et al.
Publicado: (2026)
por: Bang, Jehyeon, et al.
Publicado: (2026)
Beyond Trusting Trust: Multi-Model Validation for Robust Code Generation
por: McDanel, Bradley
Publicado: (2025)
por: McDanel, Bradley
Publicado: (2025)
PipeSpec: Breaking Stage Dependencies in Hierarchical LLM Decoding
por: McDanel, Bradley, et al.
Publicado: (2025)
por: McDanel, Bradley, et al.
Publicado: (2025)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
por: Wang, Wenfeng, et al.
Publicado: (2025)
por: Wang, Wenfeng, et al.
Publicado: (2025)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
por: Liu, Baihui, et al.
Publicado: (2026)
por: Liu, Baihui, et al.
Publicado: (2026)
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
por: Huang, Zongle, et al.
Publicado: (2025)
por: Huang, Zongle, et al.
Publicado: (2025)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
por: Xiao, Zilin, et al.
Publicado: (2024)
por: Xiao, Zilin, et al.
Publicado: (2024)
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
por: Huang, Haiduo, et al.
Publicado: (2025)
por: Huang, Haiduo, et al.
Publicado: (2025)
SERE: Similarity-based Expert Re-routing for Efficient Batch Decoding in MoE Models
por: Wu, Juntong, et al.
Publicado: (2026)
por: Wu, Juntong, et al.
Publicado: (2026)
MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenarios
por: Li, Shuhuai, et al.
Publicado: (2026)
por: Li, Shuhuai, et al.
Publicado: (2026)
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
por: Xue, Leyang, et al.
Publicado: (2024)
por: Xue, Leyang, et al.
Publicado: (2024)
Speculative Decoding and Beyond: An In-Depth Survey of Techniques
por: Hu, Yunhai, et al.
Publicado: (2025)
por: Hu, Yunhai, et al.
Publicado: (2025)
JacQuant: STE-Free Quantization-Aware Training via Learned Jacobian Surrogates
por: Yi, Kai, et al.
Publicado: (2026)
por: Yi, Kai, et al.
Publicado: (2026)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
por: Li, Lujun, et al.
Publicado: (2025)
por: Li, Lujun, et al.
Publicado: (2025)
EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU Utilization
por: Wu, Yize, et al.
Publicado: (2025)
por: Wu, Yize, et al.
Publicado: (2025)
AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference
por: Zhong, Shuzhang, et al.
Publicado: (2024)
por: Zhong, Shuzhang, et al.
Publicado: (2024)
DySpec: Faster Speculative Decoding with Dynamic Token Tree Structure
por: Xiong, Yunfan, et al.
Publicado: (2024)
por: Xiong, Yunfan, et al.
Publicado: (2024)
HiSpec: Hierarchical Speculative Decoding for LLMs
por: Kumar, Avinash, et al.
Publicado: (2025)
por: Kumar, Avinash, et al.
Publicado: (2025)
MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
por: Miao, Ruijie, et al.
Publicado: (2025)
por: Miao, Ruijie, et al.
Publicado: (2025)
Horseshoe Mixtures-of-Experts (HS-MoE)
por: Polson, Nick, et al.
Publicado: (2026)
por: Polson, Nick, et al.
Publicado: (2026)
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
por: Chen, Xiaodong, et al.
Publicado: (2025)
por: Chen, Xiaodong, et al.
Publicado: (2025)
Expert Routing for Communication-Efficient MoE via Finite Expert Banks
por: Salehi, Mohammad Reza Deylam, et al.
Publicado: (2026)
por: Salehi, Mohammad Reza Deylam, et al.
Publicado: (2026)
TriSpec: Ternary Speculative Decoding via Lightweight Proxy Verification
por: Jiang, Haoyun, et al.
Publicado: (2026)
por: Jiang, Haoyun, et al.
Publicado: (2026)
R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
por: Zhang, Zhenyu, et al.
Publicado: (2025)
por: Zhang, Zhenyu, et al.
Publicado: (2025)
SpecMemo: Speculative Decoding is in Your Pocket
por: Yildirim, Selin, et al.
Publicado: (2025)
por: Yildirim, Selin, et al.
Publicado: (2025)
SpecPipe: Accelerating Pipeline Parallelism-based LLM Inference with Speculative Decoding
por: Yin, Haofei, et al.
Publicado: (2025)
por: Yin, Haofei, et al.
Publicado: (2025)
SpecMD: A Comprehensive Study On Speculative Expert Prefetching
por: Hoang, Duc, et al.
Publicado: (2026)
por: Hoang, Duc, et al.
Publicado: (2026)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
por: Jin, Peng, et al.
Publicado: (2024)
por: Jin, Peng, et al.
Publicado: (2024)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
por: Yang, Penghui, et al.
Publicado: (2025)
por: Yang, Penghui, et al.
Publicado: (2025)
MoE Lens -- An Expert Is All You Need
por: Chaudhari, Marmik, et al.
Publicado: (2026)
por: Chaudhari, Marmik, et al.
Publicado: (2026)
MoE Pathfinder: Trajectory-driven Expert Pruning
por: Yang, Xican, et al.
Publicado: (2025)
por: Yang, Xican, et al.
Publicado: (2025)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
por: Hou, Yunlong, et al.
Publicado: (2025)
por: Hou, Yunlong, et al.
Publicado: (2025)
SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding
por: Li, Shenggui, et al.
Publicado: (2026)
por: Li, Shenggui, et al.
Publicado: (2026)
VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading
por: Xu, Cheng, et al.
Publicado: (2026)
por: Xu, Cheng, et al.
Publicado: (2026)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
por: Takashiro, Shota, et al.
Publicado: (2026)
por: Takashiro, Shota, et al.
Publicado: (2026)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
por: Liu, Xinyi, et al.
Publicado: (2026)
por: Liu, Xinyi, et al.
Publicado: (2026)
DyMoE: Dynamic Expert Orchestration with Mixed-Precision Quantization for Efficient MoE Inference on Edge
por: Huang, Yuegui, et al.
Publicado: (2026)
por: Huang, Yuegui, et al.
Publicado: (2026)
Ejemplares similares
-
CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
por: McDanel, Bradley, et al.
Publicado: (2026) -
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
por: McDanel, Bradley
Publicado: (2024) -
DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
por: Yu, Fengze, et al.
Publicado: (2025) -
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding
por: Bang, Jehyeon, et al.
Publicado: (2026) -
Beyond Trusting Trust: Multi-Model Validation for Robust Code Generation
por: McDanel, Bradley
Publicado: (2025)