From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
Fuente:
arXiv
Guardado en:
| Autores principales: | Lee, Gunjun, Kim, Jiwon, Park, Jaiyoung, Lee, Younjoo, Ahn, Jung Ho |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-Layer Scheduling for MoE-Based LLM Reasoning
por: Sun, Yifan, et al.
Publicado: (2026)
por: Sun, Yifan, et al.
Publicado: (2026)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
por: Woo, Sunghyeon, et al.
Publicado: (2026)
por: Woo, Sunghyeon, et al.
Publicado: (2026)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
por: Wang, Haodong, et al.
Publicado: (2025)
por: Wang, Haodong, et al.
Publicado: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
por: Hu, Jianmin, et al.
Publicado: (2025)
por: Hu, Jianmin, et al.
Publicado: (2025)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
por: Yu, Yanpeng, et al.
Publicado: (2025)
por: Yu, Yanpeng, et al.
Publicado: (2025)
A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
por: Zhang, Yue, et al.
Publicado: (2025)
por: Zhang, Yue, et al.
Publicado: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
por: Qian, Yulei, et al.
Publicado: (2024)
por: Qian, Yulei, et al.
Publicado: (2024)
FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving
por: Hsieh, Chia-chi, et al.
Publicado: (2026)
por: Hsieh, Chia-chi, et al.
Publicado: (2026)
MixServe: An Automatic Distributed Serving System for MoE Models with Hybrid Parallelism Based on Fused Communication Algorithm
por: Zhou, Bowen, et al.
Publicado: (2026)
por: Zhou, Bowen, et al.
Publicado: (2026)
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
por: Zhao, Lu, et al.
Publicado: (2025)
por: Zhao, Lu, et al.
Publicado: (2025)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
por: Liu, Ziming, et al.
Publicado: (2025)
por: Liu, Ziming, et al.
Publicado: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
por: Zhong, Yinmin, et al.
Publicado: (2024)
por: Zhong, Yinmin, et al.
Publicado: (2024)
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
por: Yuan, Yichao, et al.
Publicado: (2025)
por: Yuan, Yichao, et al.
Publicado: (2025)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
por: Zhang, Yuning, et al.
Publicado: (2025)
por: Zhang, Yuning, et al.
Publicado: (2025)
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
por: Li, Yan, et al.
Publicado: (2025)
por: Li, Yan, et al.
Publicado: (2025)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
por: Nie, Xiaonan, et al.
Publicado: (2024)
por: Nie, Xiaonan, et al.
Publicado: (2024)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
por: Chen, Liangkun, et al.
Publicado: (2025)
por: Chen, Liangkun, et al.
Publicado: (2025)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
por: Wang, Liujianfu, et al.
Publicado: (2025)
por: Wang, Liujianfu, et al.
Publicado: (2025)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
por: Han, Yu, et al.
Publicado: (2025)
por: Han, Yu, et al.
Publicado: (2025)
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling
por: Yang, Yuchen, et al.
Publicado: (2026)
por: Yang, Yuchen, et al.
Publicado: (2026)
ProMoE: Fast MoE-based LLM Serving using Proactive Caching
por: Song, Xiaoniu, et al.
Publicado: (2024)
por: Song, Xiaoniu, et al.
Publicado: (2024)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
por: Wang, Chao, et al.
Publicado: (2025)
por: Wang, Chao, et al.
Publicado: (2025)
LAPS: A Length-Aware-Prefill LLM Serving System
por: She, Jianshu, et al.
Publicado: (2026)
por: She, Jianshu, et al.
Publicado: (2026)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
por: Huang, En-Ming, et al.
Publicado: (2025)
por: Huang, En-Ming, et al.
Publicado: (2025)
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
por: Go, Seokjin, et al.
Publicado: (2026)
por: Go, Seokjin, et al.
Publicado: (2026)
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
por: Shi, Long, et al.
Publicado: (2025)
por: Shi, Long, et al.
Publicado: (2025)
FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training
por: Qi, Shuyao, et al.
Publicado: (2026)
por: Qi, Shuyao, et al.
Publicado: (2026)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
por: Chen, Xing, et al.
Publicado: (2025)
por: Chen, Xing, et al.
Publicado: (2025)
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
por: Lin, Wenxiang, et al.
Publicado: (2025)
por: Lin, Wenxiang, et al.
Publicado: (2025)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
por: Zhong, Shuzhang, et al.
Publicado: (2025)
por: Zhong, Shuzhang, et al.
Publicado: (2025)
Accelerating Distributed MoE Training and Inference with Lina
por: Li, Jiamin, et al.
Publicado: (2022)
por: Li, Jiamin, et al.
Publicado: (2022)
Sparse Checkpointing for Fast and Reliable MoE Training
por: Gandhi, Swapnil, et al.
Publicado: (2024)
por: Gandhi, Swapnil, et al.
Publicado: (2024)
HarMoEny: Efficient Multi-GPU Inference of MoE Models
por: Doucet, Zachary, et al.
Publicado: (2025)
por: Doucet, Zachary, et al.
Publicado: (2025)
Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference
por: Siavashi, Mohammad, et al.
Publicado: (2025)
por: Siavashi, Mohammad, et al.
Publicado: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
por: Pan, Xinglin, et al.
Publicado: (2025)
por: Pan, Xinglin, et al.
Publicado: (2025)
Fine-grained MoE Load Balancing with Linear Programming
por: Zhao, Chenqi, et al.
Publicado: (2025)
por: Zhao, Chenqi, et al.
Publicado: (2025)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
por: Luo, Jiajun, et al.
Publicado: (2024)
por: Luo, Jiajun, et al.
Publicado: (2024)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
por: Zhang, Zhexiang, et al.
Publicado: (2025)
por: Zhang, Zhexiang, et al.
Publicado: (2025)
Making MoE-based LLM Inference Resilient with Tarragon
por: Zhang, Songyu, et al.
Publicado: (2026)
por: Zhang, Songyu, et al.
Publicado: (2026)
MoEless: Efficient MoE LLM Serving via Serverless Computing
por: Yu, Hanfei, et al.
Publicado: (2026)
por: Yu, Hanfei, et al.
Publicado: (2026)
Ejemplares similares
-
Multi-Layer Scheduling for MoE-Based LLM Reasoning
por: Sun, Yifan, et al.
Publicado: (2026) -
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
por: Woo, Sunghyeon, et al.
Publicado: (2026) -
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
por: Wang, Haodong, et al.
Publicado: (2025) -
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
por: Hu, Jianmin, et al.
Publicado: (2025) -
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
por: Yu, Yanpeng, et al.
Publicado: (2025)