Multi-level Memory-Centric Profiling on ARM Processors with ARM SPE
Fuente:
arXiv
Saved in:
| Main Authors: | Miksits, Samuel, Shi, Ruimin, Gokhale, Maya, Wahlgren, Jacob, Schieffer, Gabin, Peng, Ivy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
by: Shi, Ruimin, et al.
Published: (2026)
by: Shi, Ruimin, et al.
Published: (2026)
ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
by: Shi, Ruimin, et al.
Published: (2025)
by: Shi, Ruimin, et al.
Published: (2025)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
by: Wahlgren, Jacob, et al.
Published: (2024)
by: Wahlgren, Jacob, et al.
Published: (2024)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
by: Wahlgren, Jacob, et al.
Published: (2025)
by: Wahlgren, Jacob, et al.
Published: (2025)
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
by: Schieffer, Gabin, et al.
Published: (2025)
by: Schieffer, Gabin, et al.
Published: (2025)
Understanding Layered Portability from HPC to Cloud in Containerized Environments
by: Medeiros, Daniel, et al.
Published: (2024)
by: Medeiros, Daniel, et al.
Published: (2024)
A GPU-accelerated Molecular Docking Workflow with Kubernetes and Apache Airflow
by: Medeiros, Daniel, et al.
Published: (2024)
by: Medeiros, Daniel, et al.
Published: (2024)
Kub: Enabling Elastic HPC Workloads on Containerized Environments
by: Medeiros, Daniel, et al.
Published: (2024)
by: Medeiros, Daniel, et al.
Published: (2024)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
by: Schieffer, Gabin, et al.
Published: (2026)
by: Schieffer, Gabin, et al.
Published: (2026)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
by: Shi, Ruimin, et al.
Published: (2026)
by: Shi, Ruimin, et al.
Published: (2026)
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
Communication Offloading on SmartNIC DPUs: A Quantitative Approach
by: Wahlgren, Jacob, et al.
Published: (2026)
by: Wahlgren, Jacob, et al.
Published: (2026)
Harnessing CUDA-Q's MPS for Tensor Network Simulations of Large-Scale Quantum Circuits
by: Schieffer, Gabin, et al.
Published: (2025)
by: Schieffer, Gabin, et al.
Published: (2025)
OpenCUBE: Building an Open Source Cloud Blueprint with EPI Systems
by: Peng, Ivy, et al.
Published: (2024)
by: Peng, Ivy, et al.
Published: (2024)
A New Family of Thread to Core Allocation Policies for an SMT ARM Processor
by: Navarro, Marta, et al.
Published: (2025)
by: Navarro, Marta, et al.
Published: (2025)
Exploring the Viability of Unikernels for ARM-powered Edge Computing
by: Kaiser, Shahidullah, et al.
Published: (2024)
by: Kaiser, Shahidullah, et al.
Published: (2024)
Demystifying ARM SME to Optimize General Matrix Multiplications
by: Deng, Chencheng, et al.
Published: (2025)
by: Deng, Chencheng, et al.
Published: (2025)
Computational Performance and Energy Efficiency of ARM based HPC servers
by: Schirmer, Oskar
Published: (2024)
by: Schirmer, Oskar
Published: (2024)
ARC-V: Vertical Resource Adaptivity for HPC Workloads in Containerized Environments
by: Medeiros, Daniel, et al.
Published: (2025)
by: Medeiros, Daniel, et al.
Published: (2025)
Comparison of Vectorization Capabilities of Different Compilers for X86 and ARM CPUs
by: Sakib, Nazmus, et al.
Published: (2025)
by: Sakib, Nazmus, et al.
Published: (2025)
A Hybrid Vectorized Merge Sort on ARM NEON
by: Zhou, Jincheng, et al.
Published: (2024)
by: Zhou, Jincheng, et al.
Published: (2024)
On the Performance of Cloud-based ARM SVE for Zero-Knowledge Proving Systems
by: Loghin, Dumitrel, et al.
Published: (2025)
by: Loghin, Dumitrel, et al.
Published: (2025)
oneDAL Optimization for ARM Scalable Vector Extension: Maximizing Efficiency for High-Performance Data Science
by: Sharma, Chandan, et al.
Published: (2025)
by: Sharma, Chandan, et al.
Published: (2025)
Large Scale Finite-Temperature Real-time Time Dependent Density Functional Theory Calculation with Hybrid Functional on ARM and GPU Systems
by: Liu, Rongrong, et al.
Published: (2025)
by: Liu, Rongrong, et al.
Published: (2025)
Leveraging HPC Profiling & Tracing Tools to Understand the Performance of Particle-in-Cell Monte Carlo Simulations
by: Williams, Jeremy J., et al.
Published: (2023)
by: Williams, Jeremy J., et al.
Published: (2023)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
by: Ekelund, Jonah, et al.
Published: (2025)
by: Ekelund, Jonah, et al.
Published: (2025)
An Edge-Computing based Industrial Gateway for Industry 4.0 using ARM TrustZone Technology
by: Gupta, Sandeep
Published: (2024)
by: Gupta, Sandeep
Published: (2024)
CkIO: Parallel File Input for Over-Decomposed Task-Based Systems
by: Jacob, Mathew, et al.
Published: (2024)
by: Jacob, Mathew, et al.
Published: (2024)
ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
DawnPiper: A Memory-scablable Pipeline Parallel Training Framework
by: Peng, Xuan, et al.
Published: (2025)
by: Peng, Xuan, et al.
Published: (2025)
Memory-Centric Computing: Solving Computing's Memory Problem
by: Mutlu, Onur, et al.
Published: (2025)
by: Mutlu, Onur, et al.
Published: (2025)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
by: Qianli, Liu, et al.
Published: (2025)
by: Qianli, Liu, et al.
Published: (2025)
Energy-Aware Scheduling Strategies for Partially-Replicable Task Chains on Heterogeneous Processors
by: Idouar, Yacine, et al.
Published: (2025)
by: Idouar, Yacine, et al.
Published: (2025)
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
by: Qiang, Xinwei, et al.
Published: (2026)
by: Qiang, Xinwei, et al.
Published: (2026)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
by: Zhuang, Chen, et al.
Published: (2024)
by: Zhuang, Chen, et al.
Published: (2024)
System-Level Performance Modeling of Photonic In-Memory Computing
by: Arockiaraj, Jebacyril, et al.
Published: (2026)
by: Arockiaraj, Jebacyril, et al.
Published: (2026)
Memory-Centric Computing: Recent Advances in Processing-in-DRAM
by: Mutlu, Onur, et al.
Published: (2024)
by: Mutlu, Onur, et al.
Published: (2024)
PATSMA: Parameter Auto-tuning for Shared Memory Algorithms
by: Fernandes, Joao B., et al.
Published: (2024)
by: Fernandes, Joao B., et al.
Published: (2024)
Similar Items
-
High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
by: Shi, Ruimin, et al.
Published: (2026) -
ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
by: Shi, Ruimin, et al.
Published: (2025) -
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
by: Wahlgren, Jacob, et al.
Published: (2024) -
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
by: Wahlgren, Jacob, et al.
Published: (2025) -
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
by: Schieffer, Gabin, et al.
Published: (2025)