MixServe: An Automatic Distributed Serving System for MoE Models with Hybrid Parallelism Based on Fused Communication Algorithm
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Bowen, Jia, Jinrui, He, Wenhao, Zhang, Yong, Dong, Fang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
di: Liu, Ziming, et al.
Pubblicazione: (2025)
di: Liu, Ziming, et al.
Pubblicazione: (2025)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
di: Zhang, Yuning, et al.
Pubblicazione: (2025)
di: Zhang, Yuning, et al.
Pubblicazione: (2025)
Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism
di: Wu, Zhiyuan, et al.
Pubblicazione: (2025)
di: Wu, Zhiyuan, et al.
Pubblicazione: (2025)
MoLink: Distributed and Efficient Serving Framework for Large Models
di: Jin, Lewei, et al.
Pubblicazione: (2025)
di: Jin, Lewei, et al.
Pubblicazione: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
di: Wang, Haodong, et al.
Pubblicazione: (2025)
di: Wang, Haodong, et al.
Pubblicazione: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
di: Zhang, Yuning, et al.
Pubblicazione: (2026)
di: Zhang, Yuning, et al.
Pubblicazione: (2026)
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
di: Yuan, Yichao, et al.
Pubblicazione: (2025)
di: Yuan, Yichao, et al.
Pubblicazione: (2025)
MoEless: Efficient MoE LLM Serving via Serverless Computing
di: Yu, Hanfei, et al.
Pubblicazione: (2026)
di: Yu, Hanfei, et al.
Pubblicazione: (2026)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
di: Srivatsa, Vikranth, et al.
Pubblicazione: (2026)
di: Srivatsa, Vikranth, et al.
Pubblicazione: (2026)
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
di: Go, Seokjin, et al.
Pubblicazione: (2026)
di: Go, Seokjin, et al.
Pubblicazione: (2026)
ProMoE: Fast MoE-based LLM Serving using Proactive Caching
di: Song, Xiaoniu, et al.
Pubblicazione: (2024)
di: Song, Xiaoniu, et al.
Pubblicazione: (2024)
FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving
di: Gao, Shouwei, et al.
Pubblicazione: (2026)
di: Gao, Shouwei, et al.
Pubblicazione: (2026)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
di: Yu, Yanpeng, et al.
Pubblicazione: (2025)
di: Yu, Yanpeng, et al.
Pubblicazione: (2025)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
di: Xia, Yifei, et al.
Pubblicazione: (2025)
di: Xia, Yifei, et al.
Pubblicazione: (2025)
DeepServe: Serverless Large Language Model Serving at Scale
di: Hu, Junhao, et al.
Pubblicazione: (2025)
di: Hu, Junhao, et al.
Pubblicazione: (2025)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
di: Nian, Sean, et al.
Pubblicazione: (2026)
di: Nian, Sean, et al.
Pubblicazione: (2026)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
di: Mo, Zizhao, et al.
Pubblicazione: (2025)
di: Mo, Zizhao, et al.
Pubblicazione: (2025)
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
di: Lee, Gunjun, et al.
Pubblicazione: (2025)
di: Lee, Gunjun, et al.
Pubblicazione: (2025)
Cloud Native System for LLM Inference Serving
di: Xu, Minxian, et al.
Pubblicazione: (2025)
di: Xu, Minxian, et al.
Pubblicazione: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)
di: He, Wenhao, et al.
Pubblicazione: (2026)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
di: Luo, Jiajun, et al.
Pubblicazione: (2024)
di: Luo, Jiajun, et al.
Pubblicazione: (2024)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
di: Wang, Liujianfu, et al.
Pubblicazione: (2025)
di: Wang, Liujianfu, et al.
Pubblicazione: (2025)
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
di: Zhang, Zheng, et al.
Pubblicazione: (2025)
di: Zhang, Zheng, et al.
Pubblicazione: (2025)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
di: Nie, Xiaonan, et al.
Pubblicazione: (2024)
di: Nie, Xiaonan, et al.
Pubblicazione: (2024)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
di: Cao, Jiahe, et al.
Pubblicazione: (2026)
di: Cao, Jiahe, et al.
Pubblicazione: (2026)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
di: Han, Yu, et al.
Pubblicazione: (2025)
di: Han, Yu, et al.
Pubblicazione: (2025)
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling
di: Yang, Yuchen, et al.
Pubblicazione: (2026)
di: Yang, Yuchen, et al.
Pubblicazione: (2026)
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
di: Lin, Yi-Chien, et al.
Pubblicazione: (2024)
di: Lin, Yi-Chien, et al.
Pubblicazione: (2024)
Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
di: Li, Cong, et al.
Pubblicazione: (2025)
di: Li, Cong, et al.
Pubblicazione: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
di: Duan, Jiangfei, et al.
Pubblicazione: (2024)
di: Duan, Jiangfei, et al.
Pubblicazione: (2024)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
di: Wang, Shaoyu, et al.
Pubblicazione: (2025)
di: Wang, Shaoyu, et al.
Pubblicazione: (2025)
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
di: Bu, Tianci, et al.
Pubblicazione: (2026)
di: Bu, Tianci, et al.
Pubblicazione: (2026)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
Serving Compound Inference Systems on Datacenter GPUs
di: Devata, Sriram, et al.
Pubblicazione: (2026)
di: Devata, Sriram, et al.
Pubblicazione: (2026)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
MoDM: Efficient Serving for Image Generation via Mixture-of-Diffusion Models
di: Xia, Yuchen, et al.
Pubblicazione: (2025)
di: Xia, Yuchen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
di: Hu, Jianmin, et al.
Pubblicazione: (2025) -
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
di: Liu, Ziming, et al.
Pubblicazione: (2025) -
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
di: Zhang, Yuning, et al.
Pubblicazione: (2025) -
Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism
di: Wu, Zhiyuan, et al.
Pubblicazione: (2025) -
MoLink: Distributed and Efficient Serving Framework for Large Models
di: Jin, Lewei, et al.
Pubblicazione: (2025)