CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhao, Adrian, Cai, Zhenkun, Song, Zhenyu, Yu, Lingfan, Fan, Haozheng, Wu, Jun, Wang, Yida, Vijaykumar, Nandita |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
par: Wang, Shaoyu, et autres
Publié: (2025)
par: Wang, Shaoyu, et autres
Publié: (2025)
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
par: Shi, Ge, et autres
Publié: (2025)
par: Shi, Ge, et autres
Publié: (2025)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
par: Hankendi, Can, et autres
Publié: (2026)
par: Hankendi, Can, et autres
Publié: (2026)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
par: Zhu, Ruidong, et autres
Publié: (2025)
par: Zhu, Ruidong, et autres
Publié: (2025)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
par: Go, Seokjin, et autres
Publié: (2025)
par: Go, Seokjin, et autres
Publié: (2025)
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
par: Imani, HamidReza, et autres
Publié: (2025)
par: Imani, HamidReza, et autres
Publié: (2025)
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
par: Du, Zhixu, et autres
Publié: (2023)
par: Du, Zhixu, et autres
Publié: (2023)
ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
par: Singh, Gursimran, et autres
Publié: (2025)
par: Singh, Gursimran, et autres
Publié: (2025)
Elastic Mixture of Rank-Wise Experts for Knowledge Reuse in Federated Fine-Tuning
par: Wu, Yebo, et autres
Publié: (2025)
par: Wu, Yebo, et autres
Publié: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
par: Lin, Haoran, et autres
Publié: (2025)
par: Lin, Haoran, et autres
Publié: (2025)
Stateful Large Language Model Serving with Pensieve
par: Yu, Lingfan, et autres
Publié: (2023)
par: Yu, Lingfan, et autres
Publié: (2023)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
par: Luo, Shuqing, et autres
Publié: (2024)
par: Luo, Shuqing, et autres
Publié: (2024)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
par: Liu, Ziming, et autres
Publié: (2025)
par: Liu, Ziming, et autres
Publié: (2025)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
par: Wang, Zhibin, et autres
Publié: (2025)
par: Wang, Zhibin, et autres
Publié: (2025)
Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping
par: Jiang, Chenyu, et autres
Publié: (2024)
par: Jiang, Chenyu, et autres
Publié: (2024)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
par: Pan, Xinglin, et autres
Publié: (2025)
par: Pan, Xinglin, et autres
Publié: (2025)
FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving
par: Wang, Minghe, et autres
Publié: (2026)
par: Wang, Minghe, et autres
Publié: (2026)
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
par: Cai, Weilin, et autres
Publié: (2024)
par: Cai, Weilin, et autres
Publié: (2024)
FedMoE-DA: Federated Mixture of Experts via Domain Aware Fine-grained Aggregation
par: Zhan, Ziwei, et autres
Publié: (2024)
par: Zhan, Ziwei, et autres
Publié: (2024)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
par: Liu, Xinyi, et autres
Publié: (2026)
par: Liu, Xinyi, et autres
Publié: (2026)
Optimal Expert Selection for Distributed Mixture-of-Experts at the Wireless Edge
par: Qin, Shengling, et autres
Publié: (2025)
par: Qin, Shengling, et autres
Publié: (2025)
Scattered Mixture-of-Experts Implementation
par: Tan, Shawn, et autres
Publié: (2024)
par: Tan, Shawn, et autres
Publié: (2024)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
par: Cai, Weilin, et autres
Publié: (2024)
par: Cai, Weilin, et autres
Publié: (2024)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
par: Chen, Wenyan, et autres
Publié: (2026)
par: Chen, Wenyan, et autres
Publié: (2026)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
par: Jiang, Yinsicheng, et autres
Publié: (2025)
par: Jiang, Yinsicheng, et autres
Publié: (2025)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
par: Jiang, Yinsicheng, et autres
Publié: (2024)
par: Jiang, Yinsicheng, et autres
Publié: (2024)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
par: Zhang, Yuning, et autres
Publié: (2025)
par: Zhang, Yuning, et autres
Publié: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
par: Qian, Yulei, et autres
Publié: (2024)
par: Qian, Yulei, et autres
Publié: (2024)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
par: Yu, Yanpeng, et autres
Publié: (2025)
par: Yu, Yanpeng, et autres
Publié: (2025)
MoDM: Efficient Serving for Image Generation via Mixture-of-Diffusion Models
par: Xia, Yuchen, et autres
Publié: (2025)
par: Xia, Yuchen, et autres
Publié: (2025)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
par: Ahmad, Sohaib, et autres
Publié: (2024)
par: Ahmad, Sohaib, et autres
Publié: (2024)
FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
par: Gao, Yunqi, et autres
Publié: (2025)
par: Gao, Yunqi, et autres
Publié: (2025)
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
par: Skiadopoulos, Athinagoras, et autres
Publié: (2025)
par: Skiadopoulos, Athinagoras, et autres
Publié: (2025)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
par: Wu, Yongji, et autres
Publié: (2025)
par: Wu, Yongji, et autres
Publié: (2025)
Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and Splitting
par: Zhao, Zhixin, et autres
Publié: (2024)
par: Zhao, Zhixin, et autres
Publié: (2024)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
par: Liu, Di, et autres
Publié: (2026)
par: Liu, Di, et autres
Publié: (2026)
DPQuant: Efficient and Differentially-Private Model Training via Dynamic Quantization Scheduling
par: Gao, Yubo, et autres
Publié: (2025)
par: Gao, Yubo, et autres
Publié: (2025)
Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
par: Yu, Hanfei, et autres
Publié: (2025)
par: Yu, Hanfei, et autres
Publié: (2025)
The Cost of Garbage Collection for State Machine Replication
par: Liang, Zhiying, et autres
Publié: (2024)
par: Liang, Zhiying, et autres
Publié: (2024)
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
par: Zhang, Shulai, et autres
Publié: (2025)
par: Zhang, Shulai, et autres
Publié: (2025)
Documents similaires
-
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
par: Wang, Shaoyu, et autres
Publié: (2025) -
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
par: Shi, Ge, et autres
Publié: (2025) -
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
par: Hankendi, Can, et autres
Publié: (2026) -
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
par: Zhu, Ruidong, et autres
Publié: (2025) -
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
par: Go, Seokjin, et autres
Publié: (2025)