DeepCEE: Efficient Cross-Region Model Distributed Training System under Heterogeneous GPUs and Networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Jinquan, Liao, Xiaojian, Liu, Xuzhao, Suo, Jiashun, Huo, Zhisheng, Zhang, Chenhao, Xu, Xiangrong, Shen, Runnan, Xie, Xilong, Xiao, Limin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
von: Suo, Jiashun, et al.
Veröffentlicht: (2025)
von: Suo, Jiashun, et al.
Veröffentlicht: (2025)
PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
von: Liu, Chongpeng, et al.
Veröffentlicht: (2025)
von: Liu, Chongpeng, et al.
Veröffentlicht: (2025)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
von: Sun, Mo, et al.
Veröffentlicht: (2024)
von: Sun, Mo, et al.
Veröffentlicht: (2024)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
An Adaptive Distributed Stencil Abstraction for GPUs
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
von: He, Xuan, et al.
Veröffentlicht: (2025)
von: He, Xuan, et al.
Veröffentlicht: (2025)
CB-SpMV:A Data Aggregating and Balance Algorithm for Cache-Friendly Block-Based SpMV on GPUs
von: Cong, Xing, et al.
Veröffentlicht: (2026)
von: Cong, Xing, et al.
Veröffentlicht: (2026)
Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain
von: Jang, Insu, et al.
Veröffentlicht: (2026)
von: Jang, Insu, et al.
Veröffentlicht: (2026)
Heta: Distributed Training of Heterogeneous Graph Neural Networks
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
von: Yan, Ran, et al.
Veröffentlicht: (2025)
von: Yan, Ran, et al.
Veröffentlicht: (2025)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
von: Mehboob, Talha, et al.
Veröffentlicht: (2025)
von: Mehboob, Talha, et al.
Veröffentlicht: (2025)
Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
von: Strati, Foteini, et al.
Veröffentlicht: (2025)
von: Strati, Foteini, et al.
Veröffentlicht: (2025)
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
Distributed Retrieval-Augmented Generation
von: Xu, Chenhao, et al.
Veröffentlicht: (2025)
von: Xu, Chenhao, et al.
Veröffentlicht: (2025)
Nezha: Breaking Multi-Rail Network Barriers for Distributed DNN Training
von: Yu, Enda, et al.
Veröffentlicht: (2024)
von: Yu, Enda, et al.
Veröffentlicht: (2024)
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
von: He, Guoliang, et al.
Veröffentlicht: (2025)
von: He, Guoliang, et al.
Veröffentlicht: (2025)
GOGH: Correlation-Guided Orchestration of GPUs in Heterogeneous Clusters
von: Raeisi, Ahmad, et al.
Veröffentlicht: (2025)
von: Raeisi, Ahmad, et al.
Veröffentlicht: (2025)
Accelerating Maximal Biclique Enumeration on GPUs
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
Parallelizing Maximal Clique Enumeration on GPUs
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
Optimizing sDTW for AMD GPUs
von: Latta-Lin, Daniel, et al.
Veröffentlicht: (2024)
von: Latta-Lin, Daniel, et al.
Veröffentlicht: (2024)
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
von: Chen, Kefu, et al.
Veröffentlicht: (2026)
von: Chen, Kefu, et al.
Veröffentlicht: (2026)
Cross-region Model Training with Communication-Computation Overlapping and Delay Compensation
von: Zhu, Ying, et al.
Veröffentlicht: (2025)
von: Zhu, Ying, et al.
Veröffentlicht: (2025)
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
Serving Compound Inference Systems on Datacenter GPUs
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
Optimal Workload Placement on Multi-Instance GPUs
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
Accurate Computation of the Logarithm of Modified Bessel Functions on GPUs
von: Plesner, Andreas, et al.
Veröffentlicht: (2024)
von: Plesner, Andreas, et al.
Veröffentlicht: (2024)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
von: Xu, Yaodan, et al.
Veröffentlicht: (2025)
von: Xu, Yaodan, et al.
Veröffentlicht: (2025)
Managing Multi Instance GPUs for High Throughput and Energy Savings
von: Saraha, Abhijeet, et al.
Veröffentlicht: (2025)
von: Saraha, Abhijeet, et al.
Veröffentlicht: (2025)
Analytical Performance Estimation during Code Generation on Modern GPUs
von: Ernst, Dominik, et al.
Veröffentlicht: (2022)
von: Ernst, Dominik, et al.
Veröffentlicht: (2022)
Astra: Efficient and Money-saving Automatic Parallel Strategies Search on Heterogeneous GPUs
von: Wang, Peiran, et al.
Veröffentlicht: (2025)
von: Wang, Peiran, et al.
Veröffentlicht: (2025)
Joint Training on AMD and NVIDIA GPUs
von: Hu, Jon, et al.
Veröffentlicht: (2026)
von: Hu, Jon, et al.
Veröffentlicht: (2026)
Training LLMs with Fault Tolerant HSDP on 100,000 GPUs
von: Salpekar, Omkar, et al.
Veröffentlicht: (2026)
von: Salpekar, Omkar, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
von: Suo, Jiashun, et al.
Veröffentlicht: (2025) -
PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
von: Liu, Chongpeng, et al.
Veröffentlicht: (2025) -
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
von: Sun, Mo, et al.
Veröffentlicht: (2024) -
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
von: Kim, Heehoon, et al.
Veröffentlicht: (2026) -
An Adaptive Distributed Stencil Abstraction for GPUs
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)