DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pati, Suchita, Aga, Shaizeen, Islam, Mahzabeen, Quach, Ryan, Kudchadker, Saleel, Ibrahim, Mohamed Assem
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917398778478592
author Pati, Suchita
Aga, Shaizeen
Islam, Mahzabeen
Quach, Ryan
Kudchadker, Saleel
Ibrahim, Mohamed Assem
author_facet Pati, Suchita
Aga, Shaizeen
Islam, Mahzabeen
Quach, Ryan
Kudchadker, Saleel
Ibrahim, Mohamed Assem
contents Offloading communication to existing direct memory access (DMA) engines, available on most state-of-the-art commercial GPUs, has emerged as an interesting and low-cost solution to efficiently overlap computation and communication in machine learning (ML). That said, so far, the reach of DMA offloads has been limited to bandwidth-bound scenarios only (10s of MB to GB transfer sizes). In this work, we aim to break this barrier and expand the reach of DMA communication offloads to even latency-bound regions (KB to low MB). Specifically, we discuss in this work hitherto untapped features available in the state-of-the-art AMD Instinct$^{\mathrm{TM}}$ MI300X GPUs that render DMA communication offloads competitive even for latency-bound regions. We demonstrate the efficacy of these features at the operator-level (ML communication collectives such as all-gather and all-to-all), and also at the end-to-end workload-level (LLM inference). For the former, our optimized DMA offloads close up to 4.5$\times$ performance gap and deliver additional power savings (3-10%) for ML collectives as compared to state-of-the-art GPU core-based communication library, RCCL. For the latter, we demonstrate acceleration for LLM inference: up to 1.5$\times$ lower latency and up to 1.9$\times$ higher throughput over the state-of-the-art vLLM inference framework. We conclude with a discussion of AMD Instinct GPU runtime innovations that stand to expose these features and additionally identify future hardware-software co-design potential to further improve DMA offload efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06605
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
Pati, Suchita
Aga, Shaizeen
Islam, Mahzabeen
Quach, Ryan
Kudchadker, Saleel
Ibrahim, Mohamed Assem
Distributed, Parallel, and Cluster Computing
Hardware Architecture
Offloading communication to existing direct memory access (DMA) engines, available on most state-of-the-art commercial GPUs, has emerged as an interesting and low-cost solution to efficiently overlap computation and communication in machine learning (ML). That said, so far, the reach of DMA offloads has been limited to bandwidth-bound scenarios only (10s of MB to GB transfer sizes). In this work, we aim to break this barrier and expand the reach of DMA communication offloads to even latency-bound regions (KB to low MB). Specifically, we discuss in this work hitherto untapped features available in the state-of-the-art AMD Instinct$^{\mathrm{TM}}$ MI300X GPUs that render DMA communication offloads competitive even for latency-bound regions. We demonstrate the efficacy of these features at the operator-level (ML communication collectives such as all-gather and all-to-all), and also at the end-to-end workload-level (LLM inference). For the former, our optimized DMA offloads close up to 4.5$\times$ performance gap and deliver additional power savings (3-10%) for ML collectives as compared to state-of-the-art GPU core-based communication library, RCCL. For the latter, we demonstrate acceleration for LLM inference: up to 1.5$\times$ lower latency and up to 1.9$\times$ higher throughput over the state-of-the-art vLLM inference framework. We conclude with a discussion of AMD Instinct GPU runtime innovations that stand to expose these features and additionally identify future hardware-software co-design potential to further improve DMA offload efficiency.
title DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
topic Distributed, Parallel, and Cluster Computing
Hardware Architecture
url https://arxiv.org/abs/2511.06605