Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wahlgren, Jacob, Schieffer, Gabin, Shi, Ruimin, León, Edgar A., Pearce, Roger, Gokhale, Maya, Peng, Ivy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2024)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2024)
Multi-level Memory-Centric Profiling on ARM Processors with ARM SPE
von: Miksits, Samuel, et al.
Veröffentlicht: (2024)
von: Miksits, Samuel, et al.
Veröffentlicht: (2024)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
von: Schieffer, Gabin, et al.
Veröffentlicht: (2026)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2026)
A GPU-accelerated Molecular Docking Workflow with Kubernetes and Apache Airflow
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
Understanding Layered Portability from HPC to Cloud in Containerized Environments
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
Kub: Enabling Elastic HPC Workloads on Containerized Environments
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
Comparing CPU and GPU compute of PERMANOVA on MI300A
von: Sfiligoi, Igor
Veröffentlicht: (2025)
von: Sfiligoi, Igor
Veröffentlicht: (2025)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
Communication Offloading on SmartNIC DPUs: A Quantitative Approach
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2026)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2026)
ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
von: Shi, Ruimin, et al.
Veröffentlicht: (2025)
von: Shi, Ruimin, et al.
Veröffentlicht: (2025)
High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
Dissecting the software-based measurement of CPU energy consumption: a comparative analysis
von: Raffin, Guillaume, et al.
Veröffentlicht: (2024)
von: Raffin, Guillaume, et al.
Veröffentlicht: (2024)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
GPUVM: GPU-driven Unified Virtual Memory
von: Nazaraliyev, Nurlan, et al.
Veröffentlicht: (2024)
von: Nazaraliyev, Nurlan, et al.
Veröffentlicht: (2024)
Harnessing CUDA-Q's MPS for Tensor Network Simulations of Large-Scale Quantum Circuits
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
von: Tharwani, Jay, et al.
Veröffentlicht: (2025)
von: Tharwani, Jay, et al.
Veröffentlicht: (2025)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
Porting HPC Applications to AMD Instinct$^\text{TM}$ MI300A Using Unified Memory and OpenMP
von: Tandon, Suyash, et al.
Veröffentlicht: (2024)
von: Tandon, Suyash, et al.
Veröffentlicht: (2024)
Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
OpenCUBE: Building an Open Source Cloud Blueprint with EPI Systems
von: Peng, Ivy, et al.
Veröffentlicht: (2024)
von: Peng, Ivy, et al.
Veröffentlicht: (2024)
The Energy Cost of Execution-Idle in GPU Clusters
von: Lei, Yiran, et al.
Veröffentlicht: (2026)
von: Lei, Yiran, et al.
Veröffentlicht: (2026)
Scalable GPU Performance Variability Analysis framework
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
On the Partitioning of GPU Power among Multi-Instances
von: Vamja, Tirth, et al.
Veröffentlicht: (2025)
von: Vamja, Tirth, et al.
Veröffentlicht: (2025)
Disaggregated Design for GPU-Based Volumetric Data Structures
von: Meneghin, Massimiliano, et al.
Veröffentlicht: (2025)
von: Meneghin, Massimiliano, et al.
Veröffentlicht: (2025)
Taking GPU Programming Models to Task for Performance Portability
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
Profiling and optimization of multi-card GPU machine learning jobs
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
von: Pilliat, Emmanuel
Veröffentlicht: (2026)
von: Pilliat, Emmanuel
Veröffentlicht: (2026)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
von: Islam, Tanzima Z., et al.
Veröffentlicht: (2024)
von: Islam, Tanzima Z., et al.
Veröffentlicht: (2024)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025) -
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2024) -
Multi-level Memory-Centric Profiling on ARM Processors with ARM SPE
von: Miksits, Samuel, et al.
Veröffentlicht: (2024) -
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024) -
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
von: Schieffer, Gabin, et al.
Veröffentlicht: (2026)