Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
Fuente:
arXiv
Saved in:
| Main Authors: | Pal, Shagnik, Aga, Shaizeen, Pati, Suchita, Islam, Mahzabeen, John, Lizy K. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
by: Agrawal, Anirudha, et al.
Published: (2024)
by: Agrawal, Anirudha, et al.
Published: (2024)
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
by: Pati, Suchita, et al.
Published: (2025)
by: Pati, Suchita, et al.
Published: (2025)
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
by: Ibrahim, Mohamed Assem, et al.
Published: (2024)
by: Ibrahim, Mohamed Assem, et al.
Published: (2024)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
by: Pati, Suchita, et al.
Published: (2024)
by: Pati, Suchita, et al.
Published: (2024)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
by: Pati, Suchita, et al.
Published: (2024)
by: Pati, Suchita, et al.
Published: (2024)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
by: Singhania, Varsha, et al.
Published: (2024)
by: Singhania, Varsha, et al.
Published: (2024)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
by: Kurzynski, Marco, et al.
Published: (2025)
by: Kurzynski, Marco, et al.
Published: (2025)
Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
by: Kurzynski, Marco, et al.
Published: (2025)
by: Kurzynski, Marco, et al.
Published: (2025)
IOMMU Support for Virtual-Address Remote DMA in an ARMv8 environment
by: Psistakis, Antonis
Published: (2025)
by: Psistakis, Antonis
Published: (2025)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
by: Deng, Yunhao, et al.
Published: (2025)
by: Deng, Yunhao, et al.
Published: (2025)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
by: Qu, Huanyu, et al.
Published: (2025)
by: Qu, Huanyu, et al.
Published: (2025)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
by: Kong, Fanchen, et al.
Published: (2025)
by: Kong, Fanchen, et al.
Published: (2025)
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
by: Mo, Zhiwen, et al.
Published: (2026)
by: Mo, Zhiwen, et al.
Published: (2026)
SpeedMalloc: Improving Multi-threaded Applications via a Lightweight Core for Memory Allocation
by: Li, Ruihao, et al.
Published: (2025)
by: Li, Ruihao, et al.
Published: (2025)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
by: Zhou, Zhuoshan, et al.
Published: (2026)
by: Zhou, Zhuoshan, et al.
Published: (2026)
Muchisim: A Simulation Framework for Design Exploration of Multi-Chip Manycore Systems
by: Orenes-Vera, Marcelo, et al.
Published: (2023)
by: Orenes-Vera, Marcelo, et al.
Published: (2023)
DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical Modeling
by: Mei, Linyan, et al.
Published: (2022)
by: Mei, Linyan, et al.
Published: (2022)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
by: Punniyamurthy, Kishore, et al.
Published: (2023)
by: Punniyamurthy, Kishore, et al.
Published: (2023)
JExplore: Design Space Exploration Tool for Nvidia Jetson Boards
by: Kutukcu, Basar, et al.
Published: (2025)
by: Kutukcu, Basar, et al.
Published: (2025)
The DMA Streaming Framework: Kernel-Level Buffer Orchestration for High-Performance AI Data Paths
by: Graziano, Marco
Published: (2026)
by: Graziano, Marco
Published: (2026)
RapidOMS: FPGA-based Open Modification Spectral Library Searching with HD Computing
by: Pinge, Sumukh, et al.
Published: (2024)
by: Pinge, Sumukh, et al.
Published: (2024)
Memory-Centric Computing: Solving Computing's Memory Problem
by: Mutlu, Onur, et al.
Published: (2025)
by: Mutlu, Onur, et al.
Published: (2025)
Exploration of Cryptocurrency Mining-Specific GPUs in AI Applications: A Case Study of CMP 170HX
by: Kangwei, Xing
Published: (2025)
by: Kangwei, Xing
Published: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
by: McDaniel, Adam, et al.
Published: (2026)
by: McDaniel, Adam, et al.
Published: (2026)
Revisiting Computational Storage for Data Integrity and Security
by: Shi, Chao, et al.
Published: (2025)
by: Shi, Chao, et al.
Published: (2025)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
by: Sirjani, Mohammad Sadegh, et al.
Published: (2025)
by: Sirjani, Mohammad Sadegh, et al.
Published: (2025)
Memory-Centric Computing: Recent Advances in Processing-in-DRAM
by: Mutlu, Onur, et al.
Published: (2024)
by: Mutlu, Onur, et al.
Published: (2024)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
by: Zhang, Qijun, et al.
Published: (2026)
by: Zhang, Qijun, et al.
Published: (2026)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
by: Ma, Ke, et al.
Published: (2025)
by: Ma, Ke, et al.
Published: (2025)
BlockAMC: Scalable In-Memory Analog Matrix Computing for Solving Linear Systems
by: Pan, Lunshuai, et al.
Published: (2024)
by: Pan, Lunshuai, et al.
Published: (2024)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
How Fast Can Graph Computations Go on Fine-grained Parallel Architectures
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
UniFormer: Unified and Efficient Transformer for Reasoning Across General and Custom Computing
by: Ran, Zhuoheng, et al.
Published: (2025)
by: Ran, Zhuoheng, et al.
Published: (2025)
FpgaHub: Fpga-centric Hyper-heterogeneous Computing Platform for Big Data Analytics
by: Wang, Zeke, et al.
Published: (2025)
by: Wang, Zeke, et al.
Published: (2025)
Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL
by: Meyer, Marius, et al.
Published: (2024)
by: Meyer, Marius, et al.
Published: (2024)
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
by: Prakriya, Neha, et al.
Published: (2023)
by: Prakriya, Neha, et al.
Published: (2023)
Simopt-Power: Leveraging Simulation Metadata for Low-Power Design Synthesis
by: Wadhwa, Eashan, et al.
Published: (2025)
by: Wadhwa, Eashan, et al.
Published: (2025)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
by: Feng, Weigang, et al.
Published: (2025)
by: Feng, Weigang, et al.
Published: (2025)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
by: Zhang, Yichao, et al.
Published: (2026)
by: Zhang, Yichao, et al.
Published: (2026)
Similar Items
-
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
by: Agrawal, Anirudha, et al.
Published: (2024) -
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
by: Pati, Suchita, et al.
Published: (2025) -
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
by: Ibrahim, Mohamed Assem, et al.
Published: (2024) -
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
by: Pati, Suchita, et al.
Published: (2024) -
Global Optimizations & Lightweight Dynamic Logic for Concurrency
by: Pati, Suchita, et al.
Published: (2024)