ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Choi, Yuseon, Lee, Jingu, Oh, Jungjun, Whang, Sunjoo, Kim, Byeongcheol, Kim, Minsung, Yoo, Hoi-Jun, Kim, Sangjin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
by: Choi, Yuseon, et al.
Published: (2025)
by: Choi, Yuseon, et al.
Published: (2025)
SeVeDo: A Heterogeneous Transformer Accelerator for Low-Bit Inference via Hierarchical Group Quantization and SVD-Guided Mixed Precision
by: Choi, Yuseon, et al.
Published: (2025)
by: Choi, Yuseon, et al.
Published: (2025)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
by: Chen, Liangkun, et al.
Published: (2025)
by: Chen, Liangkun, et al.
Published: (2025)
MoE-Spec: Expert Budgeting for Efficient Speculative Decoding
by: McDanel, Bradley, et al.
Published: (2026)
by: McDanel, Bradley, et al.
Published: (2026)
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
by: Huang, Zongle, et al.
Published: (2025)
by: Huang, Zongle, et al.
Published: (2025)
CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution
by: Son, Muyoung, et al.
Published: (2026)
by: Son, Muyoung, et al.
Published: (2026)
MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving
by: Su, Zhaoyuan, et al.
Published: (2026)
by: Su, Zhaoyuan, et al.
Published: (2026)
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
by: Yang, Mingyu, et al.
Published: (2025)
by: Yang, Mingyu, et al.
Published: (2025)
SRVP: Strong Recollection Video Prediction Model Using Attention-Based Spatiotemporal Correlation Fusion
by: Kim, Yuseon, et al.
Published: (2025)
by: Kim, Yuseon, et al.
Published: (2025)
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding
by: Pan, Lehan, et al.
Published: (2026)
by: Pan, Lehan, et al.
Published: (2026)
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
by: Lee, Gunjun, et al.
Published: (2025)
by: Lee, Gunjun, et al.
Published: (2025)
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
by: Lee, Seokgi, et al.
Published: (2025)
by: Lee, Seokgi, et al.
Published: (2025)
LaDiMo: Layer-wise Distillation Inspired MoEfier
by: Kim, Sungyoon, et al.
Published: (2024)
by: Kim, Sungyoon, et al.
Published: (2024)
FlashMoE: Reducing SSD I/O Bottlenecks via ML-Based Cache Replacement for Mixture-of-Experts Inference on Edge Devices
by: Kim, Byeongju, et al.
Published: (2026)
by: Kim, Byeongju, et al.
Published: (2026)
MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenarios
by: Li, Shuhuai, et al.
Published: (2026)
by: Li, Shuhuai, et al.
Published: (2026)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
by: Wang, Wenfeng, et al.
Published: (2025)
by: Wang, Wenfeng, et al.
Published: (2025)
Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
by: Pan, Yue, et al.
Published: (2025)
by: Pan, Yue, et al.
Published: (2025)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
by: Wang, Haodong, et al.
Published: (2025)
by: Wang, Haodong, et al.
Published: (2025)
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
by: Yuan, Yichao, et al.
Published: (2025)
by: Yuan, Yichao, et al.
Published: (2025)
FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving
by: Liu, Qingxiu, et al.
Published: (2026)
by: Liu, Qingxiu, et al.
Published: (2026)
MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
by: Xia, Xinfeng, et al.
Published: (2025)
by: Xia, Xinfeng, et al.
Published: (2025)
MoEless: Efficient MoE LLM Serving via Serverless Computing
by: Yu, Hanfei, et al.
Published: (2026)
by: Yu, Hanfei, et al.
Published: (2026)
3D TFTs and boundary VOAs from BPS spectra of $(G,G')$ Argyres-Douglas theories
by: Kim, Minsung, et al.
Published: (2025)
by: Kim, Minsung, et al.
Published: (2025)
ProMoE: Fast MoE-based LLM Serving using Proactive Caching
by: Song, Xiaoniu, et al.
Published: (2024)
by: Song, Xiaoniu, et al.
Published: (2024)
Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding
by: Kim, Sungkyun, et al.
Published: (2025)
by: Kim, Sungkyun, et al.
Published: (2025)
Mamba Drafters for Speculative Decoding
by: Choi, Daewon, et al.
Published: (2025)
by: Choi, Daewon, et al.
Published: (2025)
FlashMoE: Fast Distributed MoE in a Single Kernel
by: Aimuyo, Osayamen Jonathan, et al.
Published: (2025)
by: Aimuyo, Osayamen Jonathan, et al.
Published: (2025)
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
by: Park, Jihoon, et al.
Published: (2025)
by: Park, Jihoon, et al.
Published: (2025)
MoE-GS: Mixture of Experts for Dynamic Gaussian Splatting
by: Jin, In-Hwan, et al.
Published: (2025)
by: Jin, In-Hwan, et al.
Published: (2025)
mHealth‐Based Self‐Management Program for Patients With Rheumatoid Arthritis: A Pilot Randomized Controlled Study
by: Hyunjin Oh, et al.
Published: (2024)
by: Hyunjin Oh, et al.
Published: (2024)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
by: Gu, Naibin, et al.
Published: (2025)
by: Gu, Naibin, et al.
Published: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
by: Hu, Jianmin, et al.
Published: (2025)
by: Hu, Jianmin, et al.
Published: (2025)
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling
by: Yang, Yuchen, et al.
Published: (2026)
by: Yang, Yuchen, et al.
Published: (2026)
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
by: Go, Seokjin, et al.
Published: (2026)
by: Go, Seokjin, et al.
Published: (2026)
GRIN: GRadient-INformed MoE
by: Liu, Liyuan, et al.
Published: (2024)
by: Liu, Liyuan, et al.
Published: (2024)
SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation
by: Li, Zichong, et al.
Published: (2025)
by: Li, Zichong, et al.
Published: (2025)
ReTAMamba: Reliability-Aware Temporal Aggregation with Mamba for Irregular Clinical Time Series Prediction
by: Kim, Jinwoong, et al.
Published: (2026)
by: Kim, Jinwoong, et al.
Published: (2026)
GroupSegment-SHAP: Shapley Value Explanations with Group-Segment Players for Multivariate Time Series
by: Kim, Jinwoong, et al.
Published: (2026)
by: Kim, Jinwoong, et al.
Published: (2026)
IKNet: Interpretable Stock Price Prediction via Keyword-Guided Integration of News and Technical Indicators
by: Kim, Jinwoong, et al.
Published: (2025)
by: Kim, Jinwoong, et al.
Published: (2025)
Similar Items
-
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
by: Choi, Yuseon, et al.
Published: (2025) -
SeVeDo: A Heterogeneous Transformer Accelerator for Low-Bit Inference via Hierarchical Group Quantization and SVD-Guided Mixed Precision
by: Choi, Yuseon, et al.
Published: (2025) -
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
by: Chen, Liangkun, et al.
Published: (2025) -
MoE-Spec: Expert Budgeting for Efficient Speculative Decoding
by: McDanel, Bradley, et al.
Published: (2026) -
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
by: Huang, Zongle, et al.
Published: (2025)