Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Junjie, Wang, Yinzhi, Liang, Xiao, Liu, Hang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
by: Liu, Hang, et al.
Published: (2025)
by: Liu, Hang, et al.
Published: (2025)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
by: Li, Junjie
Published: (2024)
by: Li, Junjie
Published: (2024)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
by: Luo, Weile, et al.
Published: (2025)
by: Luo, Weile, et al.
Published: (2025)
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
by: Zhu, Jianwei, et al.
Published: (2024)
by: Zhu, Jianwei, et al.
Published: (2024)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
Adventures with Grace Hopper AI Super Chip and the National Research Platform
by: Hurt, J. Alex, et al.
Published: (2024)
by: Hurt, J. Alex, et al.
Published: (2024)
Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips
by: Ahmed, Mahmoud, et al.
Published: (2026)
by: Ahmed, Mahmoud, et al.
Published: (2026)
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
by: Fusco, Luigi, et al.
Published: (2024)
by: Fusco, Luigi, et al.
Published: (2024)
Dissecting the NVIDIA Blackwell Architecture with Microbenchmarks
by: Jarmusch, Aaron, et al.
Published: (2025)
by: Jarmusch, Aaron, et al.
Published: (2025)
Proposal of Automatic Offloading Method in Mixed Offloading Destination Environment
by: Yamato, Yoji
Published: (2020)
by: Yamato, Yoji
Published: (2020)
A Precision Emulation Approach to the GPU Acceleration of Ab Initio Electronic Structure Calculations
by: Liu, Hang, et al.
Published: (2026)
by: Liu, Hang, et al.
Published: (2026)
Characterization-Guided GPU Fault Resilience in NVIDIA MPS
by: Liu, Rixin, et al.
Published: (2026)
by: Liu, Rixin, et al.
Published: (2026)
Nonlinear spectral clustering with C++ GraphBLAS
by: Pasadakis, Dimosthenis, et al.
Published: (2026)
by: Pasadakis, Dimosthenis, et al.
Published: (2026)
ESS: An Offload-Centric Latent-Cache Management Architecture for DeepSeek-V3.2-Exp
by: Chen, Xinhang, et al.
Published: (2025)
by: Chen, Xinhang, et al.
Published: (2025)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
by: Wahlgren, Jacob, et al.
Published: (2024)
by: Wahlgren, Jacob, et al.
Published: (2024)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
by: Ma, Chenxiang, et al.
Published: (2025)
by: Ma, Chenxiang, et al.
Published: (2025)
MT4G: A Tool for Reliable Auto-Discovery of NVIDIA and AMD GPU Compute and Memory Topologies
by: Vanecek, Stepan, et al.
Published: (2025)
by: Vanecek, Stepan, et al.
Published: (2025)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
by: Maurya, Avinash, et al.
Published: (2024)
by: Maurya, Avinash, et al.
Published: (2024)
Effective implementation of the High Performance Conjugate Gradient benchmark on GraphBLAS
by: Scolari, Alberto, et al.
Published: (2023)
by: Scolari, Alberto, et al.
Published: (2023)
Resolving Conflicts with Grace: Dynamically Concurrent Universality
by: Kuznetsov, Petr, et al.
Published: (2025)
by: Kuznetsov, Petr, et al.
Published: (2025)
Hypersparse Traffic Matrices from Suricata Network Flows using GraphBLAS
by: Houle, Michael, et al.
Published: (2024)
by: Houle, Michael, et al.
Published: (2024)
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
by: Swann, Ryan, et al.
Published: (2025)
by: Swann, Ryan, et al.
Published: (2025)
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
by: Kumar, Abhishek Vijaya, et al.
Published: (2024)
by: Kumar, Abhishek Vijaya, et al.
Published: (2024)
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
by: Hu, Zhisheng, et al.
Published: (2025)
by: Hu, Zhisheng, et al.
Published: (2025)
MemAscend: System Memory Optimization for SSD-Offloaded LLM Fine-Tuning
by: Liaw, Yong-Cheng, et al.
Published: (2025)
by: Liaw, Yong-Cheng, et al.
Published: (2025)
Distributed Massive MIMO-Aided Task Offloading in Satellite-Terrestrial Integrated Multi-Tier VEC Networks
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
SLO-Aware Task Offloading within Collaborative Vehicle Platoons
by: Sedlak, Boris, et al.
Published: (2024)
by: Sedlak, Boris, et al.
Published: (2024)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
by: Lin, Shouxu, et al.
Published: (2026)
by: Lin, Shouxu, et al.
Published: (2026)
Mapping Parallel Matrix Multiplication in GotoBLAS2 to the AMD Versal ACAP for Deep Learning
by: Lei, Jie, et al.
Published: (2024)
by: Lei, Jie, et al.
Published: (2024)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
by: Li, Zixuan, et al.
Published: (2026)
by: Li, Zixuan, et al.
Published: (2026)
GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading
by: Song, Jaeyong, et al.
Published: (2026)
by: Song, Jaeyong, et al.
Published: (2026)
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
by: Tramm, John, et al.
Published: (2024)
by: Tramm, John, et al.
Published: (2024)
To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of Accelerators
by: Ng, Nathan, et al.
Published: (2025)
by: Ng, Nathan, et al.
Published: (2025)
Performance optimization of BLAS algorithms with band matrices for RISC-V processors
by: Pirova, Anna, et al.
Published: (2025)
by: Pirova, Anna, et al.
Published: (2025)
ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
by: Shi, Ruimin, et al.
Published: (2025)
by: Shi, Ruimin, et al.
Published: (2025)
A Survey of Computation Offloading with Task Types
by: Zhang, Siqi, et al.
Published: (2023)
by: Zhang, Siqi, et al.
Published: (2023)
PIUMA: Programmable Integrated Unified Memory Architecture
by: Aananthakrishnan, Sriram, et al.
Published: (2020)
by: Aananthakrishnan, Sriram, et al.
Published: (2020)
PUSHtap: PIM-based In-Memory HTAP with Unified Data Storage Format
by: Zhao, Yilong, et al.
Published: (2025)
by: Zhao, Yilong, et al.
Published: (2025)
Computation-Bandwidth-Memory Trade-offs: A Unified Paradigm for AI Infrastructure
by: Fan, Yuankai, et al.
Published: (2025)
by: Fan, Yuankai, et al.
Published: (2025)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
by: Ding, Zhimin, et al.
Published: (2024)
by: Ding, Zhimin, et al.
Published: (2024)
Similar Items
-
A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
by: Liu, Hang, et al.
Published: (2025) -
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
by: Li, Junjie
Published: (2024) -
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
by: Luo, Weile, et al.
Published: (2025) -
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
by: Zhu, Jianwei, et al.
Published: (2024) -
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
by: Schieffer, Gabin, et al.
Published: (2024)