Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Shigang, Ben-Nun, Tal, Di Girolamo, Salvatore, Alistarh, Dan, Hoefler, Torsten |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2019
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
von: Li, Shigang, et al.
Veröffentlicht: (2020)
von: Li, Shigang, et al.
Veröffentlicht: (2020)
Inductive Loop Analysis for Practical HPC Application Optimization
von: Schaad, Philipp, et al.
Veröffentlicht: (2025)
von: Schaad, Philipp, et al.
Veröffentlicht: (2025)
Near-Optimal Sparse Allreduce for Distributed Deep Learning
von: Li, Shigang, et al.
Veröffentlicht: (2022)
von: Li, Shigang, et al.
Veröffentlicht: (2022)
Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
von: Li, Shigang, et al.
Veröffentlicht: (2021)
von: Li, Shigang, et al.
Veröffentlicht: (2021)
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
von: Khalilov, Mikhail, et al.
Veröffentlicht: (2024)
von: Khalilov, Mikhail, et al.
Veröffentlicht: (2024)
SpaDA: A Spatial Dataflow Architecture Programming Language
von: Gianinazzi, Lukas, et al.
Veröffentlicht: (2025)
von: Gianinazzi, Lukas, et al.
Veröffentlicht: (2025)
Lion Cub: Minimizing Communication Overhead in Distributed Lion
von: Ishikawa, Satoki, et al.
Veröffentlicht: (2024)
von: Ishikawa, Satoki, et al.
Veröffentlicht: (2024)
LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
von: Schultheis, Erik, et al.
Veröffentlicht: (2025)
von: Schultheis, Erik, et al.
Veröffentlicht: (2025)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
von: Chen, Chang, et al.
Veröffentlicht: (2025)
von: Chen, Chang, et al.
Veröffentlicht: (2025)
PICO: Performance Insights for Collective Operations
von: Pasqualoni, Saverio, et al.
Veröffentlicht: (2025)
von: Pasqualoni, Saverio, et al.
Veröffentlicht: (2025)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
von: Yarlagadda, Srihas, et al.
Veröffentlicht: (2025)
von: Yarlagadda, Srihas, et al.
Veröffentlicht: (2025)
Low-Depth Spatial Tree Algorithms
von: Baumann, Yves, et al.
Veröffentlicht: (2024)
von: Baumann, Yves, et al.
Veröffentlicht: (2024)
SpComm3D: A Framework for Enabling Sparse Communication in 3D Sparse Kernels
von: Abubaker, Nabil, et al.
Veröffentlicht: (2024)
von: Abubaker, Nabil, et al.
Veröffentlicht: (2024)
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
von: Chen, Tiancheng, et al.
Veröffentlicht: (2025)
von: Chen, Tiancheng, et al.
Veröffentlicht: (2025)
Simple Opinion Dynamics for No-Regret Learning
von: Lazarsfeld, John, et al.
Veröffentlicht: (2023)
von: Lazarsfeld, John, et al.
Veröffentlicht: (2023)
Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence
von: Ansaripour, Matin, et al.
Veröffentlicht: (2022)
von: Ansaripour, Matin, et al.
Veröffentlicht: (2022)
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
von: Chen, Jinfan, et al.
Veröffentlicht: (2023)
von: Chen, Jinfan, et al.
Veröffentlicht: (2023)
Chameleon: Taming Dynamic Operator Sequences for Memory-Intensive LLM Training
von: Wang, Zibo, et al.
Veröffentlicht: (2025)
von: Wang, Zibo, et al.
Veröffentlicht: (2025)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
von: Luo, Ziyue, et al.
Veröffentlicht: (2025)
von: Luo, Ziyue, et al.
Veröffentlicht: (2025)
Federated Learning with Workload Reduction through Partial Training of Client Models and Entropy-Based Data Selection
von: Shi, Hongrui, et al.
Veröffentlicht: (2024)
von: Shi, Hongrui, et al.
Veröffentlicht: (2024)
FaaSKeeper: Learning from Building Serverless Services with ZooKeeper as an Example
von: Copik, Marcin, et al.
Veröffentlicht: (2022)
von: Copik, Marcin, et al.
Veröffentlicht: (2022)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
Taming the Memory Beast: Strategies for Reliable ML Training on Kubernetes
von: Ray, Jaideep
Veröffentlicht: (2024)
von: Ray, Jaideep
Veröffentlicht: (2024)
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
von: Nichols, Daniel, et al.
Veröffentlicht: (2026)
von: Nichols, Daniel, et al.
Veröffentlicht: (2026)
Cppless: Single-Source and High-Performance Serverless Programming in C++
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
A Unifying Framework to Enable Artificial Intelligence in High Performance Computing Workflows
von: Domke, Jens, et al.
Veröffentlicht: (2025)
von: Domke, Jens, et al.
Veröffentlicht: (2025)
Accelerating Compound LLM Training Workloads with Maestro
von: Yuan, Xiulong, et al.
Veröffentlicht: (2026)
von: Yuan, Xiulong, et al.
Veröffentlicht: (2026)
Tally: Non-Intrusive Performance Isolation for Concurrent Deep Learning Workloads
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
Software Resource Disaggregation for HPC with Serverless Computing
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication
von: Gianinazzi, Lukas, et al.
Veröffentlicht: (2024)
von: Gianinazzi, Lukas, et al.
Veröffentlicht: (2024)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
Embracing Federated Learning: Enabling Weak Client Participation via Partial Model Training
von: Lee, Sunwoo, et al.
Veröffentlicht: (2024)
von: Lee, Sunwoo, et al.
Veröffentlicht: (2024)
LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
von: Won, William, et al.
Veröffentlicht: (2021)
von: Won, William, et al.
Veröffentlicht: (2021)
Asynch-SGBDT: Asynchronous Parallel Stochastic Gradient Boosting Decision Tree based on Parameters Server
von: Daning, Cheng, et al.
Veröffentlicht: (2018)
von: Daning, Cheng, et al.
Veröffentlicht: (2018)
Improving Automatic Parallel Training via Balanced Memory Workload Optimization
von: Wang, Yujie, et al.
Veröffentlicht: (2023)
von: Wang, Yujie, et al.
Veröffentlicht: (2023)
HeteroSwitch: Characterizing and Taming System-Induced Data Heterogeneity in Federated Learning
von: Kim, Gyudong, et al.
Veröffentlicht: (2024)
von: Kim, Gyudong, et al.
Veröffentlicht: (2024)
Partial Federated Learning
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
von: Fusco, Luigi, et al.
Veröffentlicht: (2024)
von: Fusco, Luigi, et al.
Veröffentlicht: (2024)
TensorSocket: Shared Data Loading for Deep Learning Training
von: Robroek, Ties, et al.
Veröffentlicht: (2024)
von: Robroek, Ties, et al.
Veröffentlicht: (2024)
ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
von: Li, Shigang, et al.
Veröffentlicht: (2020) -
Inductive Loop Analysis for Practical HPC Application Optimization
von: Schaad, Philipp, et al.
Veröffentlicht: (2025) -
Near-Optimal Sparse Allreduce for Distributed Deep Learning
von: Li, Shigang, et al.
Veröffentlicht: (2022) -
Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
von: Li, Shigang, et al.
Veröffentlicht: (2021) -
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
von: Khalilov, Mikhail, et al.
Veröffentlicht: (2024)