Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and Algorithms
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Zhiyi, Shen, Siyuan, Bonato, Tommaso, Jeaugey, Sylvain, Alexander, Cedell, Spada, Eric, Dinan, James, Hammond, Jeff, Hoefler, Torsten |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PAT: a new algorithm for all-gather and reduce-scatter operations at scale
por: Jeaugey, Sylvain
Publicado: (2025)
por: Jeaugey, Sylvain
Publicado: (2025)
ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage
por: Shen, Siyuan, et al.
Publicado: (2025)
por: Shen, Siyuan, et al.
Publicado: (2025)
SpComm3D: A Framework for Enabling Sparse Communication in 3D Sparse Kernels
por: Abubaker, Nabil, et al.
Publicado: (2024)
por: Abubaker, Nabil, et al.
Publicado: (2024)
PICO: Performance Insights for Collective Operations
por: Pasqualoni, Saverio, et al.
Publicado: (2025)
por: Pasqualoni, Saverio, et al.
Publicado: (2025)
GPU-Initiated Networking for NCCL
por: Hamidouche, Khaled, et al.
Publicado: (2025)
por: Hamidouche, Khaled, et al.
Publicado: (2025)
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
por: Chen, Tiancheng, et al.
Publicado: (2025)
por: Chen, Tiancheng, et al.
Publicado: (2025)
Cppless: Single-Source and High-Performance Serverless Programming in C++
por: Copik, Marcin, et al.
Publicado: (2024)
por: Copik, Marcin, et al.
Publicado: (2024)
FaaSKeeper: Learning from Building Serverless Services with ZooKeeper as an Example
por: Copik, Marcin, et al.
Publicado: (2022)
por: Copik, Marcin, et al.
Publicado: (2022)
Software Resource Disaggregation for HPC with Serverless Computing
por: Copik, Marcin, et al.
Publicado: (2024)
por: Copik, Marcin, et al.
Publicado: (2024)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
por: McFarland, Thomas, et al.
Publicado: (2025)
por: McFarland, Thomas, et al.
Publicado: (2025)
A Unified CPU-GPU Protocol for GNN Training
por: Lin, Yi-Chien, et al.
Publicado: (2024)
por: Lin, Yi-Chien, et al.
Publicado: (2024)
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
por: Fusco, Luigi, et al.
Publicado: (2024)
por: Fusco, Luigi, et al.
Publicado: (2024)
Inductive Loop Analysis for Practical HPC Application Optimization
por: Schaad, Philipp, et al.
Publicado: (2025)
por: Schaad, Philipp, et al.
Publicado: (2025)
High Performance Unstructured SpMM Computation Using Tensor Cores
por: Okanovic, Patrik, et al.
Publicado: (2024)
por: Okanovic, Patrik, et al.
Publicado: (2024)
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
por: Khalilov, Mikhail, et al.
Publicado: (2024)
por: Khalilov, Mikhail, et al.
Publicado: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
por: Sojoodi, Amirhossein, et al.
Publicado: (2026)
por: Sojoodi, Amirhossein, et al.
Publicado: (2026)
Iterating Pointers: Enabling Static Analysis for Loop-based Pointers
por: Lepori, Andrea, et al.
Publicado: (2025)
por: Lepori, Andrea, et al.
Publicado: (2025)
Understanding GPU Triggering APIs for MPI+X Communication
por: Bridges, Patrick G., et al.
Publicado: (2024)
por: Bridges, Patrick G., et al.
Publicado: (2024)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
por: Shen, Aofeng, et al.
Publicado: (2025)
por: Shen, Aofeng, et al.
Publicado: (2025)
GPU-Accelerated Selected Basis Diagonalization with Thrust for SQD-based Algorithms
por: Doi, Jun, et al.
Publicado: (2026)
por: Doi, Jun, et al.
Publicado: (2026)
Demystifying the Communication Characteristics for Distributed Transformer Models
por: Anthony, Quentin, et al.
Publicado: (2024)
por: Anthony, Quentin, et al.
Publicado: (2024)
Swing: Short-cutting Rings for Higher Bandwidth Allreduce
por: De Sensi, Daniele, et al.
Publicado: (2024)
por: De Sensi, Daniele, et al.
Publicado: (2024)
Demystifying ARM SME to Optimize General Matrix Multiplications
por: Deng, Chencheng, et al.
Publicado: (2025)
por: Deng, Chencheng, et al.
Publicado: (2025)
gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
por: Huang, Jiajun, et al.
Publicado: (2023)
por: Huang, Jiajun, et al.
Publicado: (2023)
FastGraph: Optimized GPU-Enabled Algorithms for Fast Graph Building and Message Passing
por: Agarwal, Aarush, et al.
Publicado: (2025)
por: Agarwal, Aarush, et al.
Publicado: (2025)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
por: Jiang, Youhe, et al.
Publicado: (2025)
por: Jiang, Youhe, et al.
Publicado: (2025)
Optimizing Communication in Byzantine Agreement Protocols with Slim-HBBFT
por: Sony, Nasit S, et al.
Publicado: (2025)
por: Sony, Nasit S, et al.
Publicado: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
por: Zhang, Haolin, et al.
Publicado: (2025)
por: Zhang, Haolin, et al.
Publicado: (2025)
XaaS Containers: Performance-Portable Representation With Source and IR Containers
por: Copik, Marcin, et al.
Publicado: (2025)
por: Copik, Marcin, et al.
Publicado: (2025)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
por: Zhang, Mingjun, et al.
Publicado: (2025)
por: Zhang, Mingjun, et al.
Publicado: (2025)
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
por: Bridges, Patrick G., et al.
Publicado: (2026)
por: Bridges, Patrick G., et al.
Publicado: (2026)
FPsPIN: An FPGA-based Open-Hardware Research Platform for Processing in the Network
por: Schneider, Timo, et al.
Publicado: (2024)
por: Schneider, Timo, et al.
Publicado: (2024)
Core Hours and Carbon Credits: Incentivizing Sustainability in HPC
por: Kamatar, Alok, et al.
Publicado: (2025)
por: Kamatar, Alok, et al.
Publicado: (2025)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
por: Chen, Chang, et al.
Publicado: (2025)
por: Chen, Chang, et al.
Publicado: (2025)
Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
por: De Sensi, Daniele, et al.
Publicado: (2025)
por: De Sensi, Daniele, et al.
Publicado: (2025)
GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems
por: Shan, Baodi, et al.
Publicado: (2026)
por: Shan, Baodi, et al.
Publicado: (2026)
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
por: Lee, Seonho, et al.
Publicado: (2025)
por: Lee, Seonho, et al.
Publicado: (2025)
Minimizing CGYRO HPC Communication Costs in Ensembles with XGYRO by Sharing the Collisional Constant Tensor Structure
por: Sfiligoi, Igor, et al.
Publicado: (2025)
por: Sfiligoi, Igor, et al.
Publicado: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
por: Zhang, Yuning, et al.
Publicado: (2026)
por: Zhang, Yuning, et al.
Publicado: (2026)
GRNND: A GPU-Parallel Relative NN-Descent Algorithm for Efficient Approximate Nearest Neighbor Graph Construction
por: Li, Xiang, et al.
Publicado: (2025)
por: Li, Xiang, et al.
Publicado: (2025)
Ejemplares similares
-
PAT: a new algorithm for all-gather and reduce-scatter operations at scale
por: Jeaugey, Sylvain
Publicado: (2025) -
ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage
por: Shen, Siyuan, et al.
Publicado: (2025) -
SpComm3D: A Framework for Enabling Sparse Communication in 3D Sparse Kernels
por: Abubaker, Nabil, et al.
Publicado: (2024) -
PICO: Performance Insights for Collective Operations
por: Pasqualoni, Saverio, et al.
Publicado: (2025) -
GPU-Initiated Networking for NCCL
por: Hamidouche, Khaled, et al.
Publicado: (2025)