Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Shigang, Ben-Nun, Tal, Nadiradze, Giorgi, Di Girolamo, Salvatore, Dryden, Nikoli, Alistarh, Dan, Hoefler, Torsten |
|---|---|
| Formato: | Preprint |
| Publicado: |
2020
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
por: Li, Shigang, et al.
Publicado: (2019)
por: Li, Shigang, et al.
Publicado: (2019)
SpaDA: A Spatial Dataflow Architecture Programming Language
por: Gianinazzi, Lukas, et al.
Publicado: (2025)
por: Gianinazzi, Lukas, et al.
Publicado: (2025)
Lion Cub: Minimizing Communication Overhead in Distributed Lion
por: Ishikawa, Satoki, et al.
Publicado: (2024)
por: Ishikawa, Satoki, et al.
Publicado: (2024)
Near-Optimal Sparse Allreduce for Distributed Deep Learning
por: Li, Shigang, et al.
Publicado: (2022)
por: Li, Shigang, et al.
Publicado: (2022)
Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
por: Li, Shigang, et al.
Publicado: (2021)
por: Li, Shigang, et al.
Publicado: (2021)
Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence
por: Ansaripour, Matin, et al.
Publicado: (2022)
por: Ansaripour, Matin, et al.
Publicado: (2022)
Inductive Loop Analysis for Practical HPC Application Optimization
por: Schaad, Philipp, et al.
Publicado: (2025)
por: Schaad, Philipp, et al.
Publicado: (2025)
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
por: Chen, Jinfan, et al.
Publicado: (2023)
por: Chen, Jinfan, et al.
Publicado: (2023)
Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication
por: Gianinazzi, Lukas, et al.
Publicado: (2024)
por: Gianinazzi, Lukas, et al.
Publicado: (2024)
Low-Depth Spatial Tree Algorithms
por: Baumann, Yves, et al.
Publicado: (2024)
por: Baumann, Yves, et al.
Publicado: (2024)
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
por: Khalilov, Mikhail, et al.
Publicado: (2024)
por: Khalilov, Mikhail, et al.
Publicado: (2024)
Semaphores Augmented with a Waiting Array
por: Dice, Dave, et al.
Publicado: (2025)
por: Dice, Dave, et al.
Publicado: (2025)
VEIL: Reading Control Flow Graphs Like Code
por: Schaad, Philipp, et al.
Publicado: (2025)
por: Schaad, Philipp, et al.
Publicado: (2025)
Parallel Self-Avoiding Walks for a Low-Autocorrelation Binary Sequences Problem
por: Bošković, Borko, et al.
Publicado: (2022)
por: Bošković, Borko, et al.
Publicado: (2022)
SpComm3D: A Framework for Enabling Sparse Communication in 3D Sparse Kernels
por: Abubaker, Nabil, et al.
Publicado: (2024)
por: Abubaker, Nabil, et al.
Publicado: (2024)
Unifying Optimization and Dynamics to Parallelize Sequential Computation: A Guide to Parallel Newton Methods for Breaking Sequential Bottlenecks
por: Gonzalez, Xavier
Publicado: (2026)
por: Gonzalez, Xavier
Publicado: (2026)
Communication-Efficient Federated Learning With Data and Client Heterogeneity
por: Zakerinia, Hossein, et al.
Publicado: (2022)
por: Zakerinia, Hossein, et al.
Publicado: (2022)
Communication-Efficient, 2D Parallel Stochastic Gradient Descent for Distributed-Memory Optimization
por: Devarakonda, Aditya, et al.
Publicado: (2025)
por: Devarakonda, Aditya, et al.
Publicado: (2025)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
por: De Sensi, Daniele, et al.
Publicado: (2024)
por: De Sensi, Daniele, et al.
Publicado: (2024)
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
por: Xu, Youxuan, et al.
Publicado: (2025)
por: Xu, Youxuan, et al.
Publicado: (2025)
Asynch-SGBDT: Asynchronous Parallel Stochastic Gradient Boosting Decision Tree based on Parameters Server
por: Daning, Cheng, et al.
Publicado: (2018)
por: Daning, Cheng, et al.
Publicado: (2018)
A Unifying Framework to Enable Artificial Intelligence in High Performance Computing Workflows
por: Domke, Jens, et al.
Publicado: (2025)
por: Domke, Jens, et al.
Publicado: (2025)
LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
por: Schultheis, Erik, et al.
Publicado: (2025)
por: Schultheis, Erik, et al.
Publicado: (2025)
Stochastic well-structured transition systems
por: Aspnes, James
Publicado: (2025)
por: Aspnes, James
Publicado: (2025)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
por: Wang, Wenyi, et al.
Publicado: (2025)
por: Wang, Wenyi, et al.
Publicado: (2025)
Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
por: De Sensi, Daniele, et al.
Publicado: (2025)
por: De Sensi, Daniele, et al.
Publicado: (2025)
pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables
por: Ferreira, João Dinis, et al.
Publicado: (2021)
por: Ferreira, João Dinis, et al.
Publicado: (2021)
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
por: Shi, Jinliang, et al.
Publicado: (2024)
por: Shi, Jinliang, et al.
Publicado: (2024)
HYDRA: Breaking the Global Ordering Barrier in Multi-BFT Consensus
por: Lyu, Hanzheng, et al.
Publicado: (2025)
por: Lyu, Hanzheng, et al.
Publicado: (2025)
Computational Power of Opaque Robots
por: Feletti, Caterina, et al.
Publicado: (2024)
por: Feletti, Caterina, et al.
Publicado: (2024)
A Formal Semantics of C with OpenMP Parallelism (Extended Version)
por: Du, Ke, et al.
Publicado: (2026)
por: Du, Ke, et al.
Publicado: (2026)
Simple Opinion Dynamics for No-Regret Learning
por: Lazarsfeld, John, et al.
Publicado: (2023)
por: Lazarsfeld, John, et al.
Publicado: (2023)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
por: Zhuang, Chen, et al.
Publicado: (2024)
por: Zhuang, Chen, et al.
Publicado: (2024)
NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
por: Goldman, Amos, et al.
Publicado: (2026)
por: Goldman, Amos, et al.
Publicado: (2026)
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
por: Chen, Tiancheng, et al.
Publicado: (2025)
por: Chen, Tiancheng, et al.
Publicado: (2025)
Cppless: Single-Source and High-Performance Serverless Programming in C++
por: Copik, Marcin, et al.
Publicado: (2024)
por: Copik, Marcin, et al.
Publicado: (2024)
Efficient Parallel Scheduling for Sparse Triangular Solvers
por: Böhnlein, Toni, et al.
Publicado: (2025)
por: Böhnlein, Toni, et al.
Publicado: (2025)
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
por: Nichols, Daniel, et al.
Publicado: (2026)
por: Nichols, Daniel, et al.
Publicado: (2026)
Iterating Pointers: Enabling Static Analysis for Loop-based Pointers
por: Lepori, Andrea, et al.
Publicado: (2025)
por: Lepori, Andrea, et al.
Publicado: (2025)
Efficiently Scheduling Parallel DAG Tasks on Identical Multiprocessors
por: Lendve, Shardul, et al.
Publicado: (2024)
por: Lendve, Shardul, et al.
Publicado: (2024)
Ejemplares similares
-
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
por: Li, Shigang, et al.
Publicado: (2019) -
SpaDA: A Spatial Dataflow Architecture Programming Language
por: Gianinazzi, Lukas, et al.
Publicado: (2025) -
Lion Cub: Minimizing Communication Overhead in Distributed Lion
por: Ishikawa, Satoki, et al.
Publicado: (2024) -
Near-Optimal Sparse Allreduce for Distributed Deep Learning
por: Li, Shigang, et al.
Publicado: (2022) -
Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
por: Li, Shigang, et al.
Publicado: (2021)