Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and Algorithms
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Zhiyi, Shen, Siyuan, Bonato, Tommaso, Jeaugey, Sylvain, Alexander, Cedell, Spada, Eric, Dinan, James, Hammond, Jeff, Hoefler, Torsten |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PAT: a new algorithm for all-gather and reduce-scatter operations at scale
by: Jeaugey, Sylvain
Published: (2025)
by: Jeaugey, Sylvain
Published: (2025)
ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage
by: Shen, Siyuan, et al.
Published: (2025)
by: Shen, Siyuan, et al.
Published: (2025)
SpComm3D: A Framework for Enabling Sparse Communication in 3D Sparse Kernels
by: Abubaker, Nabil, et al.
Published: (2024)
by: Abubaker, Nabil, et al.
Published: (2024)
PICO: Performance Insights for Collective Operations
by: Pasqualoni, Saverio, et al.
Published: (2025)
by: Pasqualoni, Saverio, et al.
Published: (2025)
GPU-Initiated Networking for NCCL
by: Hamidouche, Khaled, et al.
Published: (2025)
by: Hamidouche, Khaled, et al.
Published: (2025)
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
by: Chen, Tiancheng, et al.
Published: (2025)
by: Chen, Tiancheng, et al.
Published: (2025)
Cppless: Single-Source and High-Performance Serverless Programming in C++
by: Copik, Marcin, et al.
Published: (2024)
by: Copik, Marcin, et al.
Published: (2024)
FaaSKeeper: Learning from Building Serverless Services with ZooKeeper as an Example
by: Copik, Marcin, et al.
Published: (2022)
by: Copik, Marcin, et al.
Published: (2022)
Software Resource Disaggregation for HPC with Serverless Computing
by: Copik, Marcin, et al.
Published: (2024)
by: Copik, Marcin, et al.
Published: (2024)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
by: McFarland, Thomas, et al.
Published: (2025)
by: McFarland, Thomas, et al.
Published: (2025)
A Unified CPU-GPU Protocol for GNN Training
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
by: Fusco, Luigi, et al.
Published: (2024)
by: Fusco, Luigi, et al.
Published: (2024)
Inductive Loop Analysis for Practical HPC Application Optimization
by: Schaad, Philipp, et al.
Published: (2025)
by: Schaad, Philipp, et al.
Published: (2025)
High Performance Unstructured SpMM Computation Using Tensor Cores
by: Okanovic, Patrik, et al.
Published: (2024)
by: Okanovic, Patrik, et al.
Published: (2024)
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
by: Khalilov, Mikhail, et al.
Published: (2024)
by: Khalilov, Mikhail, et al.
Published: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
by: Sojoodi, Amirhossein, et al.
Published: (2026)
by: Sojoodi, Amirhossein, et al.
Published: (2026)
Iterating Pointers: Enabling Static Analysis for Loop-based Pointers
by: Lepori, Andrea, et al.
Published: (2025)
by: Lepori, Andrea, et al.
Published: (2025)
Understanding GPU Triggering APIs for MPI+X Communication
by: Bridges, Patrick G., et al.
Published: (2024)
by: Bridges, Patrick G., et al.
Published: (2024)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
by: Shen, Aofeng, et al.
Published: (2025)
by: Shen, Aofeng, et al.
Published: (2025)
GPU-Accelerated Selected Basis Diagonalization with Thrust for SQD-based Algorithms
by: Doi, Jun, et al.
Published: (2026)
by: Doi, Jun, et al.
Published: (2026)
Demystifying the Communication Characteristics for Distributed Transformer Models
by: Anthony, Quentin, et al.
Published: (2024)
by: Anthony, Quentin, et al.
Published: (2024)
Swing: Short-cutting Rings for Higher Bandwidth Allreduce
by: De Sensi, Daniele, et al.
Published: (2024)
by: De Sensi, Daniele, et al.
Published: (2024)
Demystifying ARM SME to Optimize General Matrix Multiplications
by: Deng, Chencheng, et al.
Published: (2025)
by: Deng, Chencheng, et al.
Published: (2025)
gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
by: Huang, Jiajun, et al.
Published: (2023)
by: Huang, Jiajun, et al.
Published: (2023)
FastGraph: Optimized GPU-Enabled Algorithms for Fast Graph Building and Message Passing
by: Agarwal, Aarush, et al.
Published: (2025)
by: Agarwal, Aarush, et al.
Published: (2025)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
Optimizing Communication in Byzantine Agreement Protocols with Slim-HBBFT
by: Sony, Nasit S, et al.
Published: (2025)
by: Sony, Nasit S, et al.
Published: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
by: Zhang, Haolin, et al.
Published: (2025)
by: Zhang, Haolin, et al.
Published: (2025)
XaaS Containers: Performance-Portable Representation With Source and IR Containers
by: Copik, Marcin, et al.
Published: (2025)
by: Copik, Marcin, et al.
Published: (2025)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
by: Zhang, Mingjun, et al.
Published: (2025)
by: Zhang, Mingjun, et al.
Published: (2025)
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
by: Bridges, Patrick G., et al.
Published: (2026)
by: Bridges, Patrick G., et al.
Published: (2026)
FPsPIN: An FPGA-based Open-Hardware Research Platform for Processing in the Network
by: Schneider, Timo, et al.
Published: (2024)
by: Schneider, Timo, et al.
Published: (2024)
Core Hours and Carbon Credits: Incentivizing Sustainability in HPC
by: Kamatar, Alok, et al.
Published: (2025)
by: Kamatar, Alok, et al.
Published: (2025)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
by: Chen, Chang, et al.
Published: (2025)
by: Chen, Chang, et al.
Published: (2025)
Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
by: De Sensi, Daniele, et al.
Published: (2025)
by: De Sensi, Daniele, et al.
Published: (2025)
GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems
by: Shan, Baodi, et al.
Published: (2026)
by: Shan, Baodi, et al.
Published: (2026)
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
by: Lee, Seonho, et al.
Published: (2025)
by: Lee, Seonho, et al.
Published: (2025)
Minimizing CGYRO HPC Communication Costs in Ensembles with XGYRO by Sharing the Collisional Constant Tensor Structure
by: Sfiligoi, Igor, et al.
Published: (2025)
by: Sfiligoi, Igor, et al.
Published: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
by: Zhang, Yuning, et al.
Published: (2026)
by: Zhang, Yuning, et al.
Published: (2026)
GRNND: A GPU-Parallel Relative NN-Descent Algorithm for Efficient Approximate Nearest Neighbor Graph Construction
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Similar Items
-
PAT: a new algorithm for all-gather and reduce-scatter operations at scale
by: Jeaugey, Sylvain
Published: (2025) -
ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage
by: Shen, Siyuan, et al.
Published: (2025) -
SpComm3D: A Framework for Enabling Sparse Communication in 3D Sparse Kernels
by: Abubaker, Nabil, et al.
Published: (2024) -
PICO: Performance Insights for Collective Operations
by: Pasqualoni, Saverio, et al.
Published: (2025) -
GPU-Initiated Networking for NCCL
by: Hamidouche, Khaled, et al.
Published: (2025)