Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
Fuente:
arXiv
Saved in:
| Main Authors: | Schieffer, Gabin, Shi, Ruimin, Markidis, Stefano, Herten, Andreas, Faj, Jennifer, Peng, Ivy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
by: Schieffer, Gabin, et al.
Published: (2025)
by: Schieffer, Gabin, et al.
Published: (2025)
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
by: Schieffer, Gabin, et al.
Published: (2026)
by: Schieffer, Gabin, et al.
Published: (2026)
Harnessing CUDA-Q's MPS for Tensor Network Simulations of Large-Scale Quantum Circuits
by: Schieffer, Gabin, et al.
Published: (2025)
by: Schieffer, Gabin, et al.
Published: (2025)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
by: Shi, Ruimin, et al.
Published: (2026)
by: Shi, Ruimin, et al.
Published: (2026)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
by: Wahlgren, Jacob, et al.
Published: (2025)
by: Wahlgren, Jacob, et al.
Published: (2025)
A GPU-accelerated Molecular Docking Workflow with Kubernetes and Apache Airflow
by: Medeiros, Daniel, et al.
Published: (2024)
by: Medeiros, Daniel, et al.
Published: (2024)
Multi-level Memory-Centric Profiling on ARM Processors with ARM SPE
by: Miksits, Samuel, et al.
Published: (2024)
by: Miksits, Samuel, et al.
Published: (2024)
Understanding Layered Portability from HPC to Cloud in Containerized Environments
by: Medeiros, Daniel, et al.
Published: (2024)
by: Medeiros, Daniel, et al.
Published: (2024)
ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
by: Shi, Ruimin, et al.
Published: (2025)
by: Shi, Ruimin, et al.
Published: (2025)
Kub: Enabling Elastic HPC Workloads on Containerized Environments
by: Medeiros, Daniel, et al.
Published: (2024)
by: Medeiros, Daniel, et al.
Published: (2024)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
by: Wahlgren, Jacob, et al.
Published: (2024)
by: Wahlgren, Jacob, et al.
Published: (2024)
OpenCUBE: Building an Open Source Cloud Blueprint with EPI Systems
by: Peng, Ivy, et al.
Published: (2024)
by: Peng, Ivy, et al.
Published: (2024)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
by: Ekelund, Jonah, et al.
Published: (2025)
by: Ekelund, Jonah, et al.
Published: (2025)
Efficient Accelerated Graph Edit Distance Computation on GPU
by: Dabah, Adel, et al.
Published: (2026)
by: Dabah, Adel, et al.
Published: (2026)
Making Room for AI: Multi-GPU Molecular Dynamics with Deep Potentials in GROMACS
by: Pennati, Luca, et al.
Published: (2026)
by: Pennati, Luca, et al.
Published: (2026)
Leveraging HPC Profiling & Tracing Tools to Understand the Performance of Particle-in-Cell Monte Carlo Simulations
by: Williams, Jeremy J., et al.
Published: (2023)
by: Williams, Jeremy J., et al.
Published: (2023)
Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
by: Shi, Ruimin, et al.
Published: (2026)
by: Shi, Ruimin, et al.
Published: (2026)
Enabling AI Deep Potentials for Ab Initio-quality Molecular Dynamics Simulations in GROMACS
by: Hu, Andong, et al.
Published: (2026)
by: Hu, Andong, et al.
Published: (2026)
What is Quantum Parallelism, Anyhow?
by: Markidis, Stefano
Published: (2024)
by: Markidis, Stefano
Published: (2024)
Characterizing the Performance of the Implicit Massively Parallel Particle-in-Cell iPIC3D Code
by: Williams, Jeremy J., et al.
Published: (2024)
by: Williams, Jeremy J., et al.
Published: (2024)
exaCB: Reproducible Continuous Benchmark Collections at Scale Leveraging an Incremental Approach
by: Badwaik, Jayesh, et al.
Published: (2026)
by: Badwaik, Jayesh, et al.
Published: (2026)
QPU Micro-Kernels for Stencil Computation
by: Markidis, Stefano, et al.
Published: (2025)
by: Markidis, Stefano, et al.
Published: (2025)
MT4G: A Tool for Reliable Auto-Discovery of NVIDIA and AMD GPU Compute and Memory Topologies
by: Vanecek, Stepan, et al.
Published: (2025)
by: Vanecek, Stepan, et al.
Published: (2025)
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
by: Fusco, Luigi, et al.
Published: (2024)
by: Fusco, Luigi, et al.
Published: (2024)
Adaptive Multidimensional Quadrature on Multi-GPU Systems
by: Tonarelli, Melanie, et al.
Published: (2025)
by: Tonarelli, Melanie, et al.
Published: (2025)
The Hitchhiker's Guide to Programming and Optimizing Cache Coherent Heterogeneous Systems: CXL, NVLink-C2C, and AMD Infinity Fabric
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
A Multi-Objective Framework for Optimizing GPU-Enabled VM Placement in Cloud Data Centers with Multi-Instance GPU Technology
by: Siavashi, Ahmad, et al.
Published: (2025)
by: Siavashi, Ahmad, et al.
Published: (2025)
Krylov Solvers for Interior Point Methods with Applications in Radiation Therapy and Support Vector Machines
by: Liu, Felix, et al.
Published: (2023)
by: Liu, Felix, et al.
Published: (2023)
Optimizing sDTW for AMD GPUs
by: Latta-Lin, Daniel, et al.
Published: (2024)
by: Latta-Lin, Daniel, et al.
Published: (2024)
Understanding the Landscape of Ampere GPU Memory Errors
by: Zhu, Zhu, et al.
Published: (2025)
by: Zhu, Zhu, et al.
Published: (2025)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
by: Wang, Tianyu, et al.
Published: (2024)
by: Wang, Tianyu, et al.
Published: (2024)
Mass Matrix Assembly on Tensor Cores for Implicit Particle-In-Cell Methods
by: Pennati, Luca, et al.
Published: (2026)
by: Pennati, Luca, et al.
Published: (2026)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
by: Andersson, Måns I., et al.
Published: (2025)
by: Andersson, Måns I., et al.
Published: (2025)
Characterizing Production GPU Workloads using System-wide Telemetry Data
by: Cankur, Onur, et al.
Published: (2025)
by: Cankur, Onur, et al.
Published: (2025)
Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems
by: Knorr, Fabian, et al.
Published: (2025)
by: Knorr, Fabian, et al.
Published: (2025)
Understanding GPU Triggering APIs for MPI+X Communication
by: Bridges, Patrick G., et al.
Published: (2024)
by: Bridges, Patrick G., et al.
Published: (2024)
Understanding GPU Resource Interference One Level Deeper
by: Elvinger, Paul, et al.
Published: (2025)
by: Elvinger, Paul, et al.
Published: (2025)
Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
by: McDaniel, Adam, et al.
Published: (2026)
by: McDaniel, Adam, et al.
Published: (2026)
Similar Items
-
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
by: Schieffer, Gabin, et al.
Published: (2024) -
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
by: Schieffer, Gabin, et al.
Published: (2025) -
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
by: Schieffer, Gabin, et al.
Published: (2026) -
Harnessing CUDA-Q's MPS for Tensor Network Simulations of Large-Scale Quantum Circuits
by: Schieffer, Gabin, et al.
Published: (2025) -
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
by: Schieffer, Gabin, et al.
Published: (2024)