MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Ruidong, Jiang, Ziheng, Jin, Chao, Wu, Peng, Stuardo, Cesar A., Wang, Dongyang, Zhang, Xinlei, Zhou, Huaping, Wei, Haoran, Cheng, Yang, Xiao, Jianzhe, Zhang, Xinyi, Liu, Lingjun, Lin, Haibin, Chang, Li-Wen, Ye, Jianxi, Yu, Xiao, Liu, Xuanzhe, Jin, Xin, Liu, Xin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
von: Jin, Chao, et al.
Veröffentlicht: (2025)
von: Jin, Chao, et al.
Veröffentlicht: (2025)
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
von: Jiang, Ziheng, et al.
Veröffentlicht: (2024)
von: Jiang, Ziheng, et al.
Veröffentlicht: (2024)
MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
von: Jin, Yibo, et al.
Veröffentlicht: (2024)
von: Jin, Yibo, et al.
Veröffentlicht: (2024)
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
von: Shi, Ge, et al.
Veröffentlicht: (2025)
von: Shi, Ge, et al.
Veröffentlicht: (2025)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
KUBEDIRECT: Unleashing the Full Power of the Cluster Manager for Serverless Computing
von: Qi, Sheng, et al.
Veröffentlicht: (2026)
von: Qi, Sheng, et al.
Veröffentlicht: (2026)
RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation
von: Jin, Chao, et al.
Veröffentlicht: (2024)
von: Jin, Chao, et al.
Veröffentlicht: (2024)
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
von: Wu, Bingyang, et al.
Veröffentlicht: (2025)
von: Wu, Bingyang, et al.
Veröffentlicht: (2025)
Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization
von: Wan, Weilin, et al.
Veröffentlicht: (2026)
von: Wan, Weilin, et al.
Veröffentlicht: (2026)
Scaling Machine Learning Interatomic Potentials with Mixtures of Experts
von: Liu, Yuzhi, et al.
Veröffentlicht: (2026)
von: Liu, Yuzhi, et al.
Veröffentlicht: (2026)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
von: Basit, Omar, et al.
Veröffentlicht: (2026)
von: Basit, Omar, et al.
Veröffentlicht: (2026)
Trinity: Disaggregating Vector Search from Prefill-Decode Disaggregation in LLM Serving
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026)
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026)
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
von: Zhang, Shulai, et al.
Veröffentlicht: (2025)
von: Zhang, Shulai, et al.
Veröffentlicht: (2025)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives
von: Zheng, Size, et al.
Veröffentlicht: (2025)
von: Zheng, Size, et al.
Veröffentlicht: (2025)
Scaling Embeddings Outperforms Scaling Experts in Language Models
von: Liu, Hong, et al.
Veröffentlicht: (2026)
von: Liu, Hong, et al.
Veröffentlicht: (2026)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
Stratified Expert Cloning for Retention-Aware Recommendation at Scale
von: Lin, Chengzhi, et al.
Veröffentlicht: (2025)
von: Lin, Chengzhi, et al.
Veröffentlicht: (2025)
MegaHan97K: A Large-Scale Dataset for Mega-Category Chinese Character Recognition with over 97K Categories
von: Zhang, Yuyi, et al.
Veröffentlicht: (2025)
von: Zhang, Yuyi, et al.
Veröffentlicht: (2025)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
von: Wang, Shaoyu, et al.
Veröffentlicht: (2025)
von: Wang, Shaoyu, et al.
Veröffentlicht: (2025)
Expert Divergence Learning for MoE-based Language Models
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
MACE: Mixture-of-Experts Accelerated Coordinate Encoding for Large-Scale Scene Localization and Rendering
von: Liu, Mingkai, et al.
Veröffentlicht: (2025)
von: Liu, Mingkai, et al.
Veröffentlicht: (2025)
Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving
von: Li, Zongze, et al.
Veröffentlicht: (2026)
von: Li, Zongze, et al.
Veröffentlicht: (2026)
Pythia: Exploiting Workflow Predictability for Efficient Agent-Native LLM Serving
von: Yu, Shan, et al.
Veröffentlicht: (2026)
von: Yu, Shan, et al.
Veröffentlicht: (2026)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
von: Chen, Xing, et al.
Veröffentlicht: (2025)
von: Chen, Xing, et al.
Veröffentlicht: (2025)
MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data
von: Jiang, Hanwen, et al.
Veröffentlicht: (2024)
von: Jiang, Hanwen, et al.
Veröffentlicht: (2024)
TennisExpert: Towards Expert-Level Analytical Sports Video Understanding
von: Liu, Zhaoyu, et al.
Veröffentlicht: (2026)
von: Liu, Zhaoyu, et al.
Veröffentlicht: (2026)
Crocodile: Cross Experts Covariance for Disentangled Learning in Multi-Domain Recommendation
von: Lin, Zhutian, et al.
Veröffentlicht: (2024)
von: Lin, Zhutian, et al.
Veröffentlicht: (2024)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
von: Jin, Chao, et al.
Veröffentlicht: (2025) -
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
von: Jiang, Ziheng, et al.
Veröffentlicht: (2024) -
MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training
von: Zhao, Juntao, et al.
Veröffentlicht: (2025) -
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024) -
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)