Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
Fuente:
arXiv
Salvato in:
| Autori principali: | Agrawal, Anirudha, Aga, Shaizeen, Pati, Suchita, Islam, Mahzabeen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
di: Pati, Suchita, et al.
Pubblicazione: (2025)
di: Pati, Suchita, et al.
Pubblicazione: (2025)
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
di: Pal, Shagnik, et al.
Pubblicazione: (2025)
di: Pal, Shagnik, et al.
Pubblicazione: (2025)
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
di: Ibrahim, Mohamed Assem, et al.
Pubblicazione: (2024)
di: Ibrahim, Mohamed Assem, et al.
Pubblicazione: (2024)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
di: Pati, Suchita, et al.
Pubblicazione: (2024)
di: Pati, Suchita, et al.
Pubblicazione: (2024)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
di: Pati, Suchita, et al.
Pubblicazione: (2024)
di: Pati, Suchita, et al.
Pubblicazione: (2024)
Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
di: Singhania, Varsha, et al.
Pubblicazione: (2024)
di: Singhania, Varsha, et al.
Pubblicazione: (2024)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
di: Punniyamurthy, Kishore, et al.
Pubblicazione: (2023)
di: Punniyamurthy, Kishore, et al.
Pubblicazione: (2023)
IOMMU Support for Virtual-Address Remote DMA in an ARMv8 environment
di: Psistakis, Antonis
Pubblicazione: (2025)
di: Psistakis, Antonis
Pubblicazione: (2025)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
di: Deng, Yunhao, et al.
Pubblicazione: (2025)
di: Deng, Yunhao, et al.
Pubblicazione: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
di: Zhang, Chen, et al.
Pubblicazione: (2026)
di: Zhang, Chen, et al.
Pubblicazione: (2026)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
di: Kong, Fanchen, et al.
Pubblicazione: (2025)
di: Kong, Fanchen, et al.
Pubblicazione: (2025)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
di: Zhou, Zhuoshan, et al.
Pubblicazione: (2026)
di: Zhou, Zhuoshan, et al.
Pubblicazione: (2026)
Enabling Mixed criticality applications for the Versal AI-Engines
di: Sprave, Vincent, et al.
Pubblicazione: (2026)
di: Sprave, Vincent, et al.
Pubblicazione: (2026)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
di: Wijeratne, Sasindu, et al.
Pubblicazione: (2024)
di: Wijeratne, Sasindu, et al.
Pubblicazione: (2024)
Datapath Combinational Equivalence Checking With Hybrid Sweeping Engines and Parallelization
di: Chen, Zhihan, et al.
Pubblicazione: (2024)
di: Chen, Zhihan, et al.
Pubblicazione: (2024)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
di: Sirjani, Mohammad Sadegh, et al.
Pubblicazione: (2025)
di: Sirjani, Mohammad Sadegh, et al.
Pubblicazione: (2025)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
di: Chung, Euijun, et al.
Pubblicazione: (2026)
di: Chung, Euijun, et al.
Pubblicazione: (2026)
HetGPU: The pursuit of making binary compatibility towards GPUs
di: Yang, Yiwei, et al.
Pubblicazione: (2025)
di: Yang, Yiwei, et al.
Pubblicazione: (2025)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
di: Elwasif, Wael, et al.
Pubblicazione: (2022)
di: Elwasif, Wael, et al.
Pubblicazione: (2022)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
di: Jarmusch, Aaron, et al.
Pubblicazione: (2026)
di: Jarmusch, Aaron, et al.
Pubblicazione: (2026)
EDEA: Efficient Dual-Engine Accelerator for Depthwise Separable Convolution with Direct Data Transfer
di: Chen, Yi, et al.
Pubblicazione: (2025)
di: Chen, Yi, et al.
Pubblicazione: (2025)
Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods
di: Fatima, Amel, et al.
Pubblicazione: (2026)
di: Fatima, Amel, et al.
Pubblicazione: (2026)
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
di: Li, Bingyao, et al.
Pubblicazione: (2024)
di: Li, Bingyao, et al.
Pubblicazione: (2024)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
di: Qiu, Tong Dong, et al.
Pubblicazione: (2023)
di: Qiu, Tong Dong, et al.
Pubblicazione: (2023)
Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU -- A Critical Review
di: Majeed, Ashiyana Abdul, et al.
Pubblicazione: (2025)
di: Majeed, Ashiyana Abdul, et al.
Pubblicazione: (2025)
Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
di: McDaniel, Adam, et al.
Pubblicazione: (2026)
di: McDaniel, Adam, et al.
Pubblicazione: (2026)
Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
di: Eleftherakis, Panagiotis-Eleftherios, et al.
Pubblicazione: (2026)
di: Eleftherakis, Panagiotis-Eleftherios, et al.
Pubblicazione: (2026)
Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL
di: Meyer, Marius, et al.
Pubblicazione: (2024)
di: Meyer, Marius, et al.
Pubblicazione: (2024)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
di: Chu, Xiaoyu, et al.
Pubblicazione: (2024)
di: Chu, Xiaoyu, et al.
Pubblicazione: (2024)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
di: Colagrande, Luca, et al.
Pubblicazione: (2026)
di: Colagrande, Luca, et al.
Pubblicazione: (2026)
Evaluation of computational and energy performance in matrix multiplication algorithms on CPU and GPU using MKL, cuBLAS and SYCL
di: Torres, L. A., et al.
Pubblicazione: (2024)
di: Torres, L. A., et al.
Pubblicazione: (2024)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
di: Kim, Hyeseong, et al.
Pubblicazione: (2026)
di: Kim, Hyeseong, et al.
Pubblicazione: (2026)
The DMA Streaming Framework: Kernel-Level Buffer Orchestration for High-Performance AI Data Paths
di: Graziano, Marco
Pubblicazione: (2026)
di: Graziano, Marco
Pubblicazione: (2026)
GigaAPI for GPU Parallelization
di: Suvarna, M., et al.
Pubblicazione: (2025)
di: Suvarna, M., et al.
Pubblicazione: (2025)
Memory-Centric Computing: Solving Computing's Memory Problem
di: Mutlu, Onur, et al.
Pubblicazione: (2025)
di: Mutlu, Onur, et al.
Pubblicazione: (2025)
GPU-Augmented OLAP Execution Engine: GPU Offloading
di: Chang, Ilsun
Pubblicazione: (2025)
di: Chang, Ilsun
Pubblicazione: (2025)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
di: Chakraborty, Abhinaba, et al.
Pubblicazione: (2025)
di: Chakraborty, Abhinaba, et al.
Pubblicazione: (2025)
Parallelizing a modern GPU simulator
di: Huerta, Rodrigo, et al.
Pubblicazione: (2025)
di: Huerta, Rodrigo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
di: Pati, Suchita, et al.
Pubblicazione: (2025) -
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
di: Pal, Shagnik, et al.
Pubblicazione: (2025) -
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
di: Ibrahim, Mohamed Assem, et al.
Pubblicazione: (2024) -
Global Optimizations & Lightweight Dynamic Logic for Concurrency
di: Pati, Suchita, et al.
Pubblicazione: (2024) -
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
di: Pati, Suchita, et al.
Pubblicazione: (2024)