Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
Fuente:
arXiv
Saved in:
| Main Authors: | Khalilov, Mikhail, Di Girolamo, Salvatore, Chrapek, Marcin, Nudelman, Rami, Bloch, Gil, Hoefler, Torsten |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
by: Fusco, Luigi, et al.
Published: (2024)
by: Fusco, Luigi, et al.
Published: (2024)
Software Resource Disaggregation for HPC with Serverless Computing
by: Copik, Marcin, et al.
Published: (2024)
by: Copik, Marcin, et al.
Published: (2024)
OSMOSIS: Enabling Multi-Tenancy in Datacenter SmartNICs
by: Khalilov, Mikhail, et al.
Published: (2023)
by: Khalilov, Mikhail, et al.
Published: (2023)
Hazel: Secure and Efficient Disaggregated Storage
by: Chrapek, Marcin, et al.
Published: (2025)
by: Chrapek, Marcin, et al.
Published: (2025)
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
by: Li, Shigang, et al.
Published: (2019)
by: Li, Shigang, et al.
Published: (2019)
Cppless: Single-Source and High-Performance Serverless Programming in C++
by: Copik, Marcin, et al.
Published: (2024)
by: Copik, Marcin, et al.
Published: (2024)
SpComm3D: A Framework for Enabling Sparse Communication in 3D Sparse Kernels
by: Abubaker, Nabil, et al.
Published: (2024)
by: Abubaker, Nabil, et al.
Published: (2024)
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
by: Chen, Tiancheng, et al.
Published: (2025)
by: Chen, Tiancheng, et al.
Published: (2025)
FaaSKeeper: Learning from Building Serverless Services with ZooKeeper as an Example
by: Copik, Marcin, et al.
Published: (2022)
by: Copik, Marcin, et al.
Published: (2022)
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
by: Chen, Jinfan, et al.
Published: (2023)
by: Chen, Jinfan, et al.
Published: (2023)
Optimal Broadcast Schedules in Logarithmic Time with Applications to Broadcast, All-Broadcast, Reduction and All-Reduction
by: Träff, Jesper Larsson
Published: (2024)
by: Träff, Jesper Larsson
Published: (2024)
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
by: Shen, Siyuan, et al.
Published: (2024)
by: Shen, Siyuan, et al.
Published: (2024)
BCM-Broadcast: A Byzantine-Tolerant Causal Broadcast Algorithm for Distributed Mobile Systems
by: NamvariTazehkand, Leila, et al.
Published: (2024)
by: NamvariTazehkand, Leila, et al.
Published: (2024)
ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage
by: Shen, Siyuan, et al.
Published: (2025)
by: Shen, Siyuan, et al.
Published: (2025)
Amortized Asynchronous Byzantine Reliable Broadcast with Optimal Resilience
by: Hu, Michael Yiqing, et al.
Published: (2026)
by: Hu, Michael Yiqing, et al.
Published: (2026)
XaaS Containers: Performance-Portable Representation With Source and IR Containers
by: Copik, Marcin, et al.
Published: (2025)
by: Copik, Marcin, et al.
Published: (2025)
Near-Optimal Communication Byzantine Reliable Broadcast under a Message Adversary
by: Albouy, Timothé, et al.
Published: (2023)
by: Albouy, Timothé, et al.
Published: (2023)
On Orchestrating Parallel Broadcasts for Distributed Ledgers
by: Sheng, Peiyao, et al.
Published: (2024)
by: Sheng, Peiyao, et al.
Published: (2024)
Workload Distribution with Rateless Encoding: A Low-Latency Computation Offloading Method within Edge Networks
by: Guo, Zhongfu, et al.
Published: (2023)
by: Guo, Zhongfu, et al.
Published: (2023)
Distributed Massive MIMO-Aided Task Offloading in Satellite-Terrestrial Integrated Multi-Tier VEC Networks
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
SeBS-Flow: Benchmarking Serverless Cloud Function Workflows
by: Schmid, Larissa, et al.
Published: (2024)
by: Schmid, Larissa, et al.
Published: (2024)
Core Hours and Carbon Credits: Incentivizing Sustainability in HPC
by: Kamatar, Alok, et al.
Published: (2025)
by: Kamatar, Alok, et al.
Published: (2025)
Near-Optimal Sparse Allreduce for Distributed Deep Learning
by: Li, Shigang, et al.
Published: (2022)
by: Li, Shigang, et al.
Published: (2022)
Bandwidth-Aware Network Topology Optimization for Decentralized Learning
by: Shen, Yipeng, et al.
Published: (2025)
by: Shen, Yipeng, et al.
Published: (2025)
DiOMP-Offloading: Toward Portable Distributed Heterogeneous OpenMP
by: Shan, Baodi, et al.
Published: (2025)
by: Shan, Baodi, et al.
Published: (2025)
Inductive Loop Analysis for Practical HPC Application Optimization
by: Schaad, Philipp, et al.
Published: (2025)
by: Schaad, Philipp, et al.
Published: (2025)
Proposal of Automatic Offloading Method in Mixed Offloading Destination Environment
by: Yamato, Yoji
Published: (2020)
by: Yamato, Yoji
Published: (2020)
High Performance Unstructured SpMM Computation Using Tensor Cores
by: Okanovic, Patrik, et al.
Published: (2024)
by: Okanovic, Patrik, et al.
Published: (2024)
XaaS: Acceleration as a Service to Enable Productive High-Performance Cloud Computing
by: Hoefler, Torsten, et al.
Published: (2024)
by: Hoefler, Torsten, et al.
Published: (2024)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
by: Lin, Shouxu, et al.
Published: (2026)
by: Lin, Shouxu, et al.
Published: (2026)
Egret: Reinforcement Mechanism for Sequential Computation Offloading in Edge Computing
by: Peng, Haosong, et al.
Published: (2024)
by: Peng, Haosong, et al.
Published: (2024)
Distributed OpenMP Offloading of OpenMC on Intel GPU MAX Accelerators
by: Fridman, Yehonatan, et al.
Published: (2024)
by: Fridman, Yehonatan, et al.
Published: (2024)
Broadcast in Almost Mixing Time
by: Paramonov, Anton, et al.
Published: (2025)
by: Paramonov, Anton, et al.
Published: (2025)
Dynamic Probabilistic Reliable Broadcast
by: Anikina, Veronika, et al.
Published: (2023)
by: Anikina, Veronika, et al.
Published: (2023)
Computation-Bandwidth-Memory Trade-offs: A Unified Paradigm for AI Infrastructure
by: Fan, Yuankai, et al.
Published: (2025)
by: Fan, Yuankai, et al.
Published: (2025)
Neuro-Inspired Task Offloading in Edge-IoT Networks Using Spiking Neural Networks
by: Rossi, Fabio Diniz
Published: (2025)
by: Rossi, Fabio Diniz
Published: (2025)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
by: Zhang, Han, et al.
Published: (2026)
by: Zhang, Han, et al.
Published: (2026)
Broadcasting on Adversarial Multiple Access Channels
by: Aldawsari, Bader A., et al.
Published: (2021)
by: Aldawsari, Bader A., et al.
Published: (2021)
Fast Byzantine Total Order Broadcast
by: Monti, Matteo, et al.
Published: (2024)
by: Monti, Matteo, et al.
Published: (2024)
To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of Accelerators
by: Ng, Nathan, et al.
Published: (2025)
by: Ng, Nathan, et al.
Published: (2025)
Similar Items
-
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
by: Fusco, Luigi, et al.
Published: (2024) -
Software Resource Disaggregation for HPC with Serverless Computing
by: Copik, Marcin, et al.
Published: (2024) -
OSMOSIS: Enabling Multi-Tenancy in Datacenter SmartNICs
by: Khalilov, Mikhail, et al.
Published: (2023) -
Hazel: Secure and Efficient Disaggregated Storage
by: Chrapek, Marcin, et al.
Published: (2025) -
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
by: Li, Shigang, et al.
Published: (2019)