DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
Fuente:
arXiv
Salvato in:
| Autori principali: | Pati, Suchita, Aga, Shaizeen, Islam, Mahzabeen, Quach, Ryan, Kudchadker, Saleel, Ibrahim, Mohamed Assem |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
di: Agrawal, Anirudha, et al.
Pubblicazione: (2024)
di: Agrawal, Anirudha, et al.
Pubblicazione: (2024)
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
di: Pal, Shagnik, et al.
Pubblicazione: (2025)
di: Pal, Shagnik, et al.
Pubblicazione: (2025)
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
di: Ibrahim, Mohamed Assem, et al.
Pubblicazione: (2024)
di: Ibrahim, Mohamed Assem, et al.
Pubblicazione: (2024)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
di: Singhania, Varsha, et al.
Pubblicazione: (2024)
di: Singhania, Varsha, et al.
Pubblicazione: (2024)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
di: Pati, Suchita, et al.
Pubblicazione: (2024)
di: Pati, Suchita, et al.
Pubblicazione: (2024)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
di: Pati, Suchita, et al.
Pubblicazione: (2024)
di: Pati, Suchita, et al.
Pubblicazione: (2024)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
di: Deng, Yunhao, et al.
Pubblicazione: (2025)
di: Deng, Yunhao, et al.
Pubblicazione: (2025)
IOMMU Support for Virtual-Address Remote DMA in an ARMv8 environment
di: Psistakis, Antonis
Pubblicazione: (2025)
di: Psistakis, Antonis
Pubblicazione: (2025)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
di: Kong, Fanchen, et al.
Pubblicazione: (2025)
di: Kong, Fanchen, et al.
Pubblicazione: (2025)
The DMA Streaming Framework: Kernel-Level Buffer Orchestration for High-Performance AI Data Paths
di: Graziano, Marco
Pubblicazione: (2026)
di: Graziano, Marco
Pubblicazione: (2026)
Optimizing Offload Performance in Heterogeneous MPSoCs
di: Colagrande, Luca, et al.
Pubblicazione: (2024)
di: Colagrande, Luca, et al.
Pubblicazione: (2024)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
di: Meng, William, et al.
Pubblicazione: (2025)
di: Meng, William, et al.
Pubblicazione: (2025)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
di: Ma, Ke, et al.
Pubblicazione: (2025)
di: Ma, Ke, et al.
Pubblicazione: (2025)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
di: Colagrande, Luca, et al.
Pubblicazione: (2025)
di: Colagrande, Luca, et al.
Pubblicazione: (2025)
Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL
di: Meyer, Marius, et al.
Pubblicazione: (2024)
di: Meyer, Marius, et al.
Pubblicazione: (2024)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
di: Punniyamurthy, Kishore, et al.
Pubblicazione: (2023)
di: Punniyamurthy, Kishore, et al.
Pubblicazione: (2023)
On-Package Memory with Universal Chiplet Interconnect Express (UCIe): A Low Power, High Bandwidth, Low Latency and Low Cost Approach
di: Sharma, Debendra Das, et al.
Pubblicazione: (2025)
di: Sharma, Debendra Das, et al.
Pubblicazione: (2025)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
di: Chu, Xiaoyu, et al.
Pubblicazione: (2024)
di: Chu, Xiaoyu, et al.
Pubblicazione: (2024)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
di: Colagrande, Luca, et al.
Pubblicazione: (2026)
di: Colagrande, Luca, et al.
Pubblicazione: (2026)
Agentic Operator Generation for ML ASICs
di: Hammond, Alec M., et al.
Pubblicazione: (2025)
di: Hammond, Alec M., et al.
Pubblicazione: (2025)
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
di: Liu, Fangxin, et al.
Pubblicazione: (2026)
di: Liu, Fangxin, et al.
Pubblicazione: (2026)
ML-QLS: Multilevel Quantum Layout Synthesis
di: Lin, Wan-Hsuan, et al.
Pubblicazione: (2024)
di: Lin, Wan-Hsuan, et al.
Pubblicazione: (2024)
Taming Asynchronous CPU-GPU Coupling for Frequency-aware Latency Estimation on Mobile Edge
di: Chen, Jiesong, et al.
Pubblicazione: (2026)
di: Chen, Jiesong, et al.
Pubblicazione: (2026)
PIUMA: Programmable Integrated Unified Memory Architecture
di: Aananthakrishnan, Sriram, et al.
Pubblicazione: (2020)
di: Aananthakrishnan, Sriram, et al.
Pubblicazione: (2020)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
di: Pan, Yudong, et al.
Pubblicazione: (2026)
di: Pan, Yudong, et al.
Pubblicazione: (2026)
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
di: DeBole, Michael V., et al.
Pubblicazione: (2025)
di: DeBole, Michael V., et al.
Pubblicazione: (2025)
MOFCO: Mobility- and Migration-Aware Task Offloading in Three-Layer Fog Computing Environments
di: Mahdizadeh, Soheil, et al.
Pubblicazione: (2025)
di: Mahdizadeh, Soheil, et al.
Pubblicazione: (2025)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
di: Kwak, Hyunseok, et al.
Pubblicazione: (2025)
di: Kwak, Hyunseok, et al.
Pubblicazione: (2025)
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
di: Kreuzer, Anke, et al.
Pubblicazione: (2019)
di: Kreuzer, Anke, et al.
Pubblicazione: (2019)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
di: Qiu, Tong Dong, et al.
Pubblicazione: (2023)
di: Qiu, Tong Dong, et al.
Pubblicazione: (2023)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
di: Xu, Weihong, et al.
Pubblicazione: (2025)
di: Xu, Weihong, et al.
Pubblicazione: (2025)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
di: Negi, Shubham, et al.
Pubblicazione: (2025)
di: Negi, Shubham, et al.
Pubblicazione: (2025)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
di: Kubo, Tatsuya, et al.
Pubblicazione: (2025)
di: Kubo, Tatsuya, et al.
Pubblicazione: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
di: Liu, Lian, et al.
Pubblicazione: (2026)
di: Liu, Lian, et al.
Pubblicazione: (2026)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
di: Chen, Yanru, et al.
Pubblicazione: (2025)
di: Chen, Yanru, et al.
Pubblicazione: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
di: Zhang, Chen, et al.
Pubblicazione: (2026)
di: Zhang, Chen, et al.
Pubblicazione: (2026)
Efficient deadlock avoidance for 2D mesh NoCs that use OQ or VOQ routers
di: Papaphilippou, Philippos, et al.
Pubblicazione: (2023)
di: Papaphilippou, Philippos, et al.
Pubblicazione: (2023)
DCRA: A Distributed Chiplet-based Reconfigurable Architecture for Irregular Applications
di: Orenes-Vera, Marcelo, et al.
Pubblicazione: (2023)
di: Orenes-Vera, Marcelo, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
di: Agrawal, Anirudha, et al.
Pubblicazione: (2024) -
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
di: Pal, Shagnik, et al.
Pubblicazione: (2025) -
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
di: Ibrahim, Mohamed Assem, et al.
Pubblicazione: (2024) -
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
di: Singhania, Varsha, et al.
Pubblicazione: (2024) -
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
di: Pati, Suchita, et al.
Pubblicazione: (2024)