UCCL-EP: Portable Expert-Parallel Communication
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Mao, Ziming, Zhang, Yihan, Cui, Chihan, Huang, Zhen, You, Kaichao, Chen, Zhongjie, Xu, Zhiying, Gu, Zhenyu, Shenker, Scott, Raiciu, Costin, Zhou, Yang, Stoica, Ion |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
UCCL-Zip: Lossless Compression Supercharged GPU Communication
par: Ma, Shuang, et autres
Publié: (2026)
par: Ma, Shuang, et autres
Publié: (2026)
Revisiting Cache Freshness for Emerging Real-Time Applications
par: Mao, Ziming, et autres
Publié: (2024)
par: Mao, Ziming, et autres
Publié: (2024)
SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
par: Xia, Tian, et autres
Publié: (2025)
par: Xia, Tian, et autres
Publié: (2025)
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
par: Mao, Ziming, et autres
Publié: (2024)
par: Mao, Ziming, et autres
Publié: (2024)
Pie: Pooling CPU Memory for LLM Inference
par: Xu, Yi, et autres
Publié: (2024)
par: Xu, Yi, et autres
Publié: (2024)
SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost
par: Li, Zhifei, et autres
Publié: (2026)
par: Li, Zhifei, et autres
Publié: (2026)
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training
par: Zheng, Size, et autres
Publié: (2026)
par: Zheng, Size, et autres
Publié: (2026)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
par: Zhang, Chen, et autres
Publié: (2025)
par: Zhang, Chen, et autres
Publié: (2025)
Delta Fair Sharing: Performance Isolation for Multi-Tenant Storage Systems
par: Griggs, Tyler, et autres
Publié: (2026)
par: Griggs, Tyler, et autres
Publié: (2026)
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP
par: Zhao, Yilong, et autres
Publié: (2026)
par: Zhao, Yilong, et autres
Publié: (2026)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
par: Wu, Yongji, et autres
Publié: (2025)
par: Wu, Yongji, et autres
Publié: (2025)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
par: Zhang, Geng, et autres
Publié: (2025)
par: Zhang, Geng, et autres
Publié: (2025)
On Optimizing the Communication of Model Parallelism
par: Zhuang, Yonghao, et autres
Publié: (2022)
par: Zhuang, Yonghao, et autres
Publié: (2022)
Parallel and (Nearly) Work-Efficient Dynamic Programming
par: Ding, Xiangyun, et autres
Publié: (2024)
par: Ding, Xiangyun, et autres
Publié: (2024)
A Parallel and Highly-Portable HPC Poisson Solver: Preconditioned Bi-CGSTAB with alpaka
par: Pennati, Luca, et autres
Publié: (2025)
par: Pennati, Luca, et autres
Publié: (2025)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
par: Zhao, Xuanlei, et autres
Publié: (2024)
par: Zhao, Xuanlei, et autres
Publié: (2024)
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
par: Wu, Yongji, et autres
Publié: (2025)
par: Wu, Yongji, et autres
Publié: (2025)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
par: Wu, Yongji, et autres
Publié: (2024)
par: Wu, Yongji, et autres
Publié: (2024)
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
par: Liu, Xueshen, et autres
Publié: (2026)
par: Liu, Xueshen, et autres
Publié: (2026)
Parallel Integer Sort: Theory and Practice
par: Dong, Xiaojun, et autres
Publié: (2024)
par: Dong, Xiaojun, et autres
Publié: (2024)
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
par: Liu, Ziming, et autres
Publié: (2024)
par: Liu, Ziming, et autres
Publié: (2024)
Portability Efficiency Approach for Calculating Performance Portability
par: Marowka, Ami
Publié: (2024)
par: Marowka, Ami
Publié: (2024)
HybridEP: Scaling Expert Parallelism to Cross-Datacenter Scenario via Hybrid Expert/Data Transmission
par: Yang, Weihao, et autres
Publié: (2025)
par: Yang, Weihao, et autres
Publié: (2025)
Parallel Cluster-BFS and Applications to Shortest Paths
par: Wang, Letong, et autres
Publié: (2024)
par: Wang, Letong, et autres
Publié: (2024)
Parallel $k$-Core Decomposition: Theory and Practice
par: Liu, Youzhe, et autres
Publié: (2025)
par: Liu, Youzhe, et autres
Publié: (2025)
PASGAL: Parallel And Scalable Graph Algorithm Library
par: Dong, Xiaojun, et autres
Publié: (2024)
par: Dong, Xiaojun, et autres
Publié: (2024)
Parallel Point-to-Point Shortest Paths and Batch Queries
par: Dong, Xiaojun, et autres
Publié: (2025)
par: Dong, Xiaojun, et autres
Publié: (2025)
Fast and Space-Efficient Parallel Algorithms for Influence Maximization
par: Wang, Letong, et autres
Publié: (2023)
par: Wang, Letong, et autres
Publié: (2023)
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
par: Chen, Jiefei, et autres
Publié: (2026)
par: Chen, Jiefei, et autres
Publié: (2026)
The Streaming Batch Model for Efficient and Fault-Tolerant Heterogeneous Execution
par: Luan, Frank Sifei, et autres
Publié: (2025)
par: Luan, Frank Sifei, et autres
Publié: (2025)
DSP: Dynamic Sequence Parallelism for Multi-Dimensional Transformers
par: Zhao, Xuanlei, et autres
Publié: (2024)
par: Zhao, Xuanlei, et autres
Publié: (2024)
HPDR: High-Performance Portable Scientific Data Reduction Framework
par: Chen, Jieyang, et autres
Publié: (2025)
par: Chen, Jieyang, et autres
Publié: (2025)
Parallel Joinable B-Trees in the Fork-Join I/O Model
par: Goodrich, Michael, et autres
Publié: (2025)
par: Goodrich, Michael, et autres
Publié: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
par: Lin, Haoran, et autres
Publié: (2025)
par: Lin, Haoran, et autres
Publié: (2025)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
par: Cai, Weilin, et autres
Publié: (2024)
par: Cai, Weilin, et autres
Publié: (2024)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
par: Duan, Jiangfei, et autres
Publié: (2024)
par: Duan, Jiangfei, et autres
Publié: (2024)
A Massively Parallel Performance Portable Free-space Spectral Poisson Solver
par: Mayani, Sonali, et autres
Publié: (2024)
par: Mayani, Sonali, et autres
Publié: (2024)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
par: Liu, Ziming, et autres
Publié: (2025)
par: Liu, Ziming, et autres
Publié: (2025)
HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUs
par: Li, Yanliang, et autres
Publié: (2025)
par: Li, Yanliang, et autres
Publié: (2025)
Parallel Contraction Hierarchies Can Be Efficient and Scalable
par: Wan, Zijin, et autres
Publié: (2024)
par: Wan, Zijin, et autres
Publié: (2024)
Documents similaires
-
UCCL-Zip: Lossless Compression Supercharged GPU Communication
par: Ma, Shuang, et autres
Publié: (2026) -
Revisiting Cache Freshness for Emerging Real-Time Applications
par: Mao, Ziming, et autres
Publié: (2024) -
SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
par: Xia, Tian, et autres
Publié: (2025) -
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
par: Mao, Ziming, et autres
Publié: (2024) -
Pie: Pooling CPU Memory for LLM Inference
par: Xu, Yi, et autres
Publié: (2024)