SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Bang, Jehyeon, Cho, Eunyeong, Hwang, Ranggi, Chung, Jinha, Rhu, Minsoo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
by: Cho, Eunyeong, et al.
Published: (2026)
by: Cho, Eunyeong, et al.
Published: (2026)
SpecMoE: Spectral Mixture-of-Experts Foundation Model for Cross-Species EEG Decoding
by: Darankoum, Davy, et al.
Published: (2026)
by: Darankoum, Davy, et al.
Published: (2026)
Agent-X: Full Pipeline Acceleration of On-device AI Agents
by: Chung, Jinha, et al.
Published: (2026)
by: Chung, Jinha, et al.
Published: (2026)
Debunking the CUDA Myth Towards GPU-based AI Systems
by: Lee, Yunjae, et al.
Published: (2024)
by: Lee, Yunjae, et al.
Published: (2024)
vTrain: A Simulation Framework for Evaluating Cost-effective and Compute-optimal Large Language Model Training
by: Bang, Jehyeon, et al.
Published: (2023)
by: Bang, Jehyeon, et al.
Published: (2023)
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
by: Hwang, Ranggi, et al.
Published: (2023)
by: Hwang, Ranggi, et al.
Published: (2023)
The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective
by: Kim, Jiin, et al.
Published: (2025)
by: Kim, Jiin, et al.
Published: (2025)
MoE-Spec: Expert Budgeting for Efficient Speculative Decoding
by: McDanel, Bradley, et al.
Published: (2026)
by: McDanel, Bradley, et al.
Published: (2026)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
by: Liu, Xing, et al.
Published: (2025)
by: Liu, Xing, et al.
Published: (2025)
SpecMD: A Comprehensive Study On Speculative Expert Prefetching
by: Hoang, Duc, et al.
Published: (2026)
by: Hoang, Duc, et al.
Published: (2026)
Speculating Experts Accelerates Inference for Mixture-of-Experts
by: Madan, Vivan, et al.
Published: (2026)
by: Madan, Vivan, et al.
Published: (2026)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
by: Ning, Zhiyuan, et al.
Published: (2025)
by: Ning, Zhiyuan, et al.
Published: (2025)
LazyDP: Co-Designing Algorithm-Software for Scalable Training of Differentially Private Recommendation Models
by: Lim, Juntaek, et al.
Published: (2024)
by: Lim, Juntaek, et al.
Published: (2024)
Utility-Driven Speculative Decoding for Mixture-of-Experts
by: Saxena, Anish, et al.
Published: (2025)
by: Saxena, Anish, et al.
Published: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
by: Tiwari, Rishabh, et al.
Published: (2025)
by: Tiwari, Rishabh, et al.
Published: (2025)
SpecExtend: A Drop-in Enhancement for Speculative Decoding of Long Sequences
by: Cha, Jungyoub, et al.
Published: (2025)
by: Cha, Jungyoub, et al.
Published: (2025)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
by: Hou, Yunlong, et al.
Published: (2025)
by: Hou, Yunlong, et al.
Published: (2025)
FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
by: Yeo, Gwangoo, et al.
Published: (2024)
by: Yeo, Gwangoo, et al.
Published: (2024)
PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models
by: Lee, Yunjae, et al.
Published: (2024)
by: Lee, Yunjae, et al.
Published: (2024)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
by: Zhou, Yongchao, et al.
Published: (2023)
by: Zhou, Yongchao, et al.
Published: (2023)
HiSpec: Hierarchical Speculative Decoding for LLMs
by: Kumar, Avinash, et al.
Published: (2025)
by: Kumar, Avinash, et al.
Published: (2025)
KnapSpec: Self-Speculative Decoding via Adaptive Layer Selection as a Knapsack Problem
by: Cha, Seongjin, et al.
Published: (2026)
by: Cha, Seongjin, et al.
Published: (2026)
SpecMemo: Speculative Decoding is in Your Pocket
by: Yildirim, Selin, et al.
Published: (2025)
by: Yildirim, Selin, et al.
Published: (2025)
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
by: Huang, Kaixuan, et al.
Published: (2024)
by: Huang, Kaixuan, et al.
Published: (2024)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
by: Yang, Penghui, et al.
Published: (2025)
by: Yang, Penghui, et al.
Published: (2025)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
by: Liu, Baihui, et al.
Published: (2026)
by: Liu, Baihui, et al.
Published: (2026)
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
by: Li, Guanghao, et al.
Published: (2025)
by: Li, Guanghao, et al.
Published: (2025)
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
by: Sun, Ryan, et al.
Published: (2024)
by: Sun, Ryan, et al.
Published: (2024)
SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding
by: Li, Shenggui, et al.
Published: (2026)
by: Li, Shenggui, et al.
Published: (2026)
EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation
by: Zhang, Shuyu, et al.
Published: (2026)
by: Zhang, Shuyu, et al.
Published: (2026)
SpecTr: Fast Speculative Decoding via Optimal Transport
by: Sun, Ziteng, et al.
Published: (2023)
by: Sun, Ziteng, et al.
Published: (2023)
Mixture of Attentions For Speculative Decoding
by: Zimmer, Matthieu, et al.
Published: (2024)
by: Zimmer, Matthieu, et al.
Published: (2024)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
by: Shen, Yuhao, et al.
Published: (2025)
by: Shen, Yuhao, et al.
Published: (2025)
MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
by: Zhao, Zhixiong, et al.
Published: (2026)
by: Zhao, Zhixiong, et al.
Published: (2026)
ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts
by: Georganas, Evangelos, et al.
Published: (2025)
by: Georganas, Evangelos, et al.
Published: (2025)
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
by: Wang, Songsheng, et al.
Published: (2025)
by: Wang, Songsheng, et al.
Published: (2025)
Dynamic-Width Speculative Beam Decoding for Efficient LLM Inference
by: Qin, Zongyue, et al.
Published: (2024)
by: Qin, Zongyue, et al.
Published: (2024)
Similar Items
-
PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
by: Cho, Eunyeong, et al.
Published: (2026) -
SpecMoE: Spectral Mixture-of-Experts Foundation Model for Cross-Species EEG Decoding
by: Darankoum, Davy, et al.
Published: (2026) -
Agent-X: Full Pipeline Acceleration of On-device AI Agents
by: Chung, Jinha, et al.
Published: (2026) -
Debunking the CUDA Myth Towards GPU-based AI Systems
by: Lee, Yunjae, et al.
Published: (2024) -
vTrain: A Simulation Framework for Evaluating Cost-effective and Compute-optimal Large Language Model Training
by: Bang, Jehyeon, et al.
Published: (2023)