Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping
Fuente:
arXiv
Guardado en:
| Autores principales: | Jiang, Chenyu, Tian, Ye, Jia, Zhen, Zheng, Shuai, Wu, Chuan, Wang, Yida |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
por: Jiang, Chenyu, et al.
Publicado: (2025)
por: Jiang, Chenyu, et al.
Publicado: (2025)
DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
por: Tian, Ye, et al.
Publicado: (2024)
por: Tian, Ye, et al.
Publicado: (2024)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
por: Xu, Guanbin, et al.
Publicado: (2026)
por: Xu, Guanbin, et al.
Publicado: (2026)
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
por: Zhang, Shulai, et al.
Publicado: (2025)
por: Zhang, Shulai, et al.
Publicado: (2025)
TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives
por: Zheng, Size, et al.
Publicado: (2025)
por: Zheng, Size, et al.
Publicado: (2025)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
por: Wang, Zhibin, et al.
Publicado: (2025)
por: Wang, Zhibin, et al.
Publicado: (2025)
Cross-region Model Training with Communication-Computation Overlapping and Delay Compensation
por: Zhu, Ying, et al.
Publicado: (2025)
por: Zhu, Ying, et al.
Publicado: (2025)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
por: Liu, Mengfan, et al.
Publicado: (2025)
por: Liu, Mengfan, et al.
Publicado: (2025)
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
por: Lee, Seonho, et al.
Publicado: (2025)
por: Lee, Seonho, et al.
Publicado: (2025)
CO2: Efficient Distributed Training with Full Communication-Computation Overlap
por: Sun, Weigao, et al.
Publicado: (2024)
por: Sun, Weigao, et al.
Publicado: (2024)
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
por: Zhao, Adrian, et al.
Publicado: (2026)
por: Zhao, Adrian, et al.
Publicado: (2026)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
por: Luo, Shuqing, et al.
Publicado: (2024)
por: Luo, Shuqing, et al.
Publicado: (2024)
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
por: Qiang, Xinwei, et al.
Publicado: (2026)
por: Qiang, Xinwei, et al.
Publicado: (2026)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
por: Jin, Chao, et al.
Publicado: (2025)
por: Jin, Chao, et al.
Publicado: (2025)
Heta: Distributed Training of Heterogeneous Graph Neural Networks
por: Zhong, Yuchen, et al.
Publicado: (2024)
por: Zhong, Yuchen, et al.
Publicado: (2024)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
por: Wu, Yongji, et al.
Publicado: (2024)
por: Wu, Yongji, et al.
Publicado: (2024)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
por: Papavasileiou, Ioannis, et al.
Publicado: (2026)
por: Papavasileiou, Ioannis, et al.
Publicado: (2026)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
por: Chen, Liangkun, et al.
Publicado: (2025)
por: Chen, Liangkun, et al.
Publicado: (2025)
FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
por: Gao, Yunqi, et al.
Publicado: (2025)
por: Gao, Yunqi, et al.
Publicado: (2025)
Communication-Computation Pipeline Parallel Split Learning over Wireless Edge Networks
por: Liu, Chenyu, et al.
Publicado: (2025)
por: Liu, Chenyu, et al.
Publicado: (2025)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
por: Cai, Weilin, et al.
Publicado: (2024)
por: Cai, Weilin, et al.
Publicado: (2024)
CondenseGraph: Communication-Efficient Distributed GNN Training via On-the-Fly Graph Condensation
por: Zhang, Zizhao, et al.
Publicado: (2026)
por: Zhang, Zizhao, et al.
Publicado: (2026)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
por: Wu, Tian, et al.
Publicado: (2025)
por: Wu, Tian, et al.
Publicado: (2025)
Elastic Mixture of Rank-Wise Experts for Knowledge Reuse in Federated Fine-Tuning
por: Wu, Yebo, et al.
Publicado: (2025)
por: Wu, Yebo, et al.
Publicado: (2025)
Efficient Accelerated Graph Edit Distance Computation on GPU
por: Dabah, Adel, et al.
Publicado: (2026)
por: Dabah, Adel, et al.
Publicado: (2026)
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
por: Shi, Long, et al.
Publicado: (2025)
por: Shi, Long, et al.
Publicado: (2025)
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
por: Skiadopoulos, Athinagoras, et al.
Publicado: (2025)
por: Skiadopoulos, Athinagoras, et al.
Publicado: (2025)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
por: Yu, Dianhai, et al.
Publicado: (2022)
por: Yu, Dianhai, et al.
Publicado: (2022)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
por: Wang, Shaoyu, et al.
Publicado: (2025)
por: Wang, Shaoyu, et al.
Publicado: (2025)
CDFGNN: a Systematic Design of Cache-based Distributed Full-Batch Graph Neural Network Training with Communication Reduction
por: Zhang, Shuai, et al.
Publicado: (2024)
por: Zhang, Shuai, et al.
Publicado: (2024)
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training
por: Zheng, Size, et al.
Publicado: (2026)
por: Zheng, Size, et al.
Publicado: (2026)
FedMoE-DA: Federated Mixture of Experts via Domain Aware Fine-grained Aggregation
por: Zhan, Ziwei, et al.
Publicado: (2024)
por: Zhan, Ziwei, et al.
Publicado: (2024)
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
por: Li, Shengwei, et al.
Publicado: (2023)
por: Li, Shengwei, et al.
Publicado: (2023)
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
por: Xu, Heng, et al.
Publicado: (2025)
por: Xu, Heng, et al.
Publicado: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
por: Lin, Haoran, et al.
Publicado: (2025)
por: Lin, Haoran, et al.
Publicado: (2025)
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
por: Zheng, Size, et al.
Publicado: (2025)
por: Zheng, Size, et al.
Publicado: (2025)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
por: Wu, Yongji, et al.
Publicado: (2025)
por: Wu, Yongji, et al.
Publicado: (2025)
Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference
por: Luo, Shuqing, et al.
Publicado: (2025)
por: Luo, Shuqing, et al.
Publicado: (2025)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
por: Liang, Yan, et al.
Publicado: (2026)
por: Liang, Yan, et al.
Publicado: (2026)
Accelerating Distributed MoE Training and Inference with Lina
por: Li, Jiamin, et al.
Publicado: (2022)
por: Li, Jiamin, et al.
Publicado: (2022)
Ejemplares similares
-
DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
por: Jiang, Chenyu, et al.
Publicado: (2025) -
DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
por: Tian, Ye, et al.
Publicado: (2024) -
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
por: Xu, Guanbin, et al.
Publicado: (2026) -
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
por: Zhang, Shulai, et al.
Publicado: (2025) -
TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives
por: Zheng, Size, et al.
Publicado: (2025)