Holmes: Towards Distributed Training Across Clusters with Heterogeneous NIC Environment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Fei, Peng, Shuang, Sun, Ning, Wang, Fangyu, Wang, Yuanyuan, Wu, Fu, Qiu, Jiezhong, Pan, Aimin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
von: Lu, Zhengxian, et al.
Veröffentlicht: (2024)
von: Lu, Zhengxian, et al.
Veröffentlicht: (2024)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
von: Strati, Foteini, et al.
Veröffentlicht: (2025)
von: Strati, Foteini, et al.
Veröffentlicht: (2025)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
von: Liang, Antian, et al.
Veröffentlicht: (2025)
von: Liang, Antian, et al.
Veröffentlicht: (2025)
Heta: Distributed Training of Heterogeneous Graph Neural Networks
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)
Towards Lock Modularization for Heterogeneous Environments
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
von: Tang, Ding, et al.
Veröffentlicht: (2025)
von: Tang, Ding, et al.
Veröffentlicht: (2025)
Communication Offloading on SmartNIC DPUs: A Quantitative Approach
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2026)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2026)
HETHUB: A Distributed Training System with Heterogeneous Cluster for Large-Scale Models
von: Xu, Si, et al.
Veröffentlicht: (2024)
von: Xu, Si, et al.
Veröffentlicht: (2024)
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
von: Chang, Zihan, et al.
Veröffentlicht: (2024)
von: Chang, Zihan, et al.
Veröffentlicht: (2024)
HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis
von: Zhang, Shiwei, et al.
Veröffentlicht: (2024)
von: Zhang, Shiwei, et al.
Veröffentlicht: (2024)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2024)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2024)
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
von: Chen, Kefu, et al.
Veröffentlicht: (2026)
von: Chen, Kefu, et al.
Veröffentlicht: (2026)
Training DNN Models over Heterogeneous Clusters with Optimal Performance
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
Cephalo: Harnessing Heterogeneous GPU Clusters for Training Transformer Models
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2024)
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2024)
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2025)
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2025)
DFPL: Decentralized Federated Prototype Learning Across Heterogeneous Data Distributions
von: Zhang, Hongliang, et al.
Veröffentlicht: (2025)
von: Zhang, Hongliang, et al.
Veröffentlicht: (2025)
Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous Platforms
von: Qiao, Tong, et al.
Veröffentlicht: (2025)
von: Qiao, Tong, et al.
Veröffentlicht: (2025)
Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain
von: Jang, Insu, et al.
Veröffentlicht: (2026)
von: Jang, Insu, et al.
Veröffentlicht: (2026)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
von: Luo, Shuqing, et al.
Veröffentlicht: (2024)
von: Luo, Shuqing, et al.
Veröffentlicht: (2024)
Toward Heterogeneous, Distributed, and Energy-Efficient Computing with SYCL
von: Cosenza, Biagio, et al.
Veröffentlicht: (2025)
von: Cosenza, Biagio, et al.
Veröffentlicht: (2025)
ColonyOS -- A Meta-Operating System for Distributed Computing Across Heterogeneous Platform
von: Kristiansson, Johan
Veröffentlicht: (2024)
von: Kristiansson, Johan
Veröffentlicht: (2024)
DiOMP-Offloading: Toward Portable Distributed Heterogeneous OpenMP
von: Shan, Baodi, et al.
Veröffentlicht: (2025)
von: Shan, Baodi, et al.
Veröffentlicht: (2025)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
von: Mehboob, Talha, et al.
Veröffentlicht: (2025)
von: Mehboob, Talha, et al.
Veröffentlicht: (2025)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
von: Warraich, Ertza, et al.
Veröffentlicht: (2025)
von: Warraich, Ertza, et al.
Veröffentlicht: (2025)
MPI-Q: A Message Communication Library for Large-Scale Classical-Quantum Heterogeneous Hybrid Distributed Computing
von: Wang, Feng, et al.
Veröffentlicht: (2026)
von: Wang, Feng, et al.
Veröffentlicht: (2026)
HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments
von: He, Yongjun, et al.
Veröffentlicht: (2025)
von: He, Yongjun, et al.
Veröffentlicht: (2025)
TrustMesh: A Blockchain-Enabled Trusted Distributed Computing Framework for Open Heterogeneous IoT Environments
von: Rangwala, Murtaza, et al.
Veröffentlicht: (2024)
von: Rangwala, Murtaza, et al.
Veröffentlicht: (2024)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
von: Yan, Ran, et al.
Veröffentlicht: (2024)
von: Yan, Ran, et al.
Veröffentlicht: (2024)
Towards Seamless Serverless Computing Across an Edge-Cloud Continuum
von: Simion, Emilian, et al.
Veröffentlicht: (2024)
von: Simion, Emilian, et al.
Veröffentlicht: (2024)
Failure-Resilient Distributed Inference with Model Compression over Heterogeneous Edge Devices
von: Wang, Li, et al.
Veröffentlicht: (2024)
von: Wang, Li, et al.
Veröffentlicht: (2024)
Dalek: An Unconventional and Energy-Aware Heterogeneous Cluster
von: Cassagne, Adrien, et al.
Veröffentlicht: (2025)
von: Cassagne, Adrien, et al.
Veröffentlicht: (2025)
COUNTER: Cluster GCN based Energy Efficient Resource Management for Sustainable Cloud Computing Environments
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
von: Liang, Yan, et al.
Veröffentlicht: (2026)
von: Liang, Yan, et al.
Veröffentlicht: (2026)
Optimal Resource Efficiency with Fairness in Heterogeneous GPU Clusters
von: Mo, Zizhao, et al.
Veröffentlicht: (2024)
von: Mo, Zizhao, et al.
Veröffentlicht: (2024)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
von: Ai, Xin, et al.
Veröffentlicht: (2024)
von: Ai, Xin, et al.
Veröffentlicht: (2024)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
von: Sun, Mo, et al.
Veröffentlicht: (2024)
von: Sun, Mo, et al.
Veröffentlicht: (2024)
The Forward-In-Time-Only Assumption in SmartNIC Resource Management: A Critique of Wave and the Case for Bilateral Interaction
von: Borrill, Paul
Veröffentlicht: (2026)
von: Borrill, Paul
Veröffentlicht: (2026)
Ähnliche Einträge
-
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
von: Lu, Zhengxian, et al.
Veröffentlicht: (2024) -
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024) -
Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
von: Strati, Foteini, et al.
Veröffentlicht: (2025) -
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
von: Liang, Antian, et al.
Veröffentlicht: (2025) -
Heta: Distributed Training of Heterogeneous Graph Neural Networks
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)