Reimagining RDMA Through the Lens of ML

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Warraich, Ertza, Imran, Ali, Zulfiqar, Annus, Vargaftik, Shay, Fahmy, Sonia, Shahbaz, Muhammad
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917048853987328
author Warraich, Ertza
Imran, Ali
Zulfiqar, Annus
Vargaftik, Shay
Fahmy, Sonia
Shahbaz, Muhammad
author_facet Warraich, Ertza
Imran, Ali
Zulfiqar, Annus
Vargaftik, Shay
Fahmy, Sonia
Shahbaz, Muhammad
contents As distributed machine learning (ML) workloads scale to thousands of GPUs connected by ultra-high-speed inter-connects, tail latency in collective communication has emerged as a primary bottleneck. Prior RDMA designs, like RoCE, IRN, and SRNIC, enforce strict reliability and in-order delivery, relying on retransmissions and packet sequencing to ensure correctness. While effective for general-purpose workloads, these mechanisms introduce complexity and latency that scale poorly, where even rare packet losses or delays can consistently degrade system performance. We introduce Celeris, a domain-specific RDMA transport that revisits traditional reliability guarantees based on ML's tolerance for lost or partial data. Celeris removes retransmissions and in-order delivery from the RDMA NIC, enabling best-effort transport that exploits the robustness of ML workloads. It retains congestion control (e.g., DCQCN) and manages communication with software-level mechanisms such as adaptive timeouts and data prioritization, while shifting loss recovery to the ML pipeline (e.g., using the Hadamard Transform). Early results show that Celeris reduces 99th-percentile latency by up to 2.3x, cuts BRAM usage by 67%, and nearly doubles NIC resilience to faults -- delivering a resilient, scalable transport tailored for ML at cluster scale.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16606
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reimagining RDMA Through the Lens of ML
Warraich, Ertza
Imran, Ali
Zulfiqar, Annus
Vargaftik, Shay
Fahmy, Sonia
Shahbaz, Muhammad
Distributed, Parallel, and Cluster Computing
Networking and Internet Architecture
As distributed machine learning (ML) workloads scale to thousands of GPUs connected by ultra-high-speed inter-connects, tail latency in collective communication has emerged as a primary bottleneck. Prior RDMA designs, like RoCE, IRN, and SRNIC, enforce strict reliability and in-order delivery, relying on retransmissions and packet sequencing to ensure correctness. While effective for general-purpose workloads, these mechanisms introduce complexity and latency that scale poorly, where even rare packet losses or delays can consistently degrade system performance. We introduce Celeris, a domain-specific RDMA transport that revisits traditional reliability guarantees based on ML's tolerance for lost or partial data. Celeris removes retransmissions and in-order delivery from the RDMA NIC, enabling best-effort transport that exploits the robustness of ML workloads. It retains congestion control (e.g., DCQCN) and manages communication with software-level mechanisms such as adaptive timeouts and data prioritization, while shifting loss recovery to the ML pipeline (e.g., using the Hadamard Transform). Early results show that Celeris reduces 99th-percentile latency by up to 2.3x, cuts BRAM usage by 67%, and nearly doubles NIC resilience to faults -- delivering a resilient, scalable transport tailored for ML at cluster scale.
title Reimagining RDMA Through the Lens of ML
topic Distributed, Parallel, and Cluster Computing
Networking and Internet Architecture
url https://arxiv.org/abs/2510.16606