ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Ruimin, Schieffer, Gabin, Gokhale, Maya, Lin, Pei-Hung, Patel, Hiren, Peng, Ivy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
Multi-level Memory-Centric Profiling on ARM Processors with ARM SPE
von: Miksits, Samuel, et al.
Veröffentlicht: (2024)
von: Miksits, Samuel, et al.
Veröffentlicht: (2024)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
von: Schieffer, Gabin, et al.
Veröffentlicht: (2026)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2026)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2024)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2024)
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025)
Kub: Enabling Elastic HPC Workloads on Containerized Environments
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
Understanding Layered Portability from HPC to Cloud in Containerized Environments
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
Harnessing CUDA-Q's MPS for Tensor Network Simulations of Large-Scale Quantum Circuits
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025)
A GPU-accelerated Molecular Docking Workflow with Kubernetes and Apache Airflow
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
On the Performance of Cloud-based ARM SVE for Zero-Knowledge Proving Systems
von: Loghin, Dumitrel, et al.
Veröffentlicht: (2025)
von: Loghin, Dumitrel, et al.
Veröffentlicht: (2025)
Computational Performance and Energy Efficiency of ARM based HPC servers
von: Schirmer, Oskar
Veröffentlicht: (2024)
von: Schirmer, Oskar
Veröffentlicht: (2024)
Communication Offloading on SmartNIC DPUs: A Quantitative Approach
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2026)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2026)
OpenCUBE: Building an Open Source Cloud Blueprint with EPI Systems
von: Peng, Ivy, et al.
Veröffentlicht: (2024)
von: Peng, Ivy, et al.
Veröffentlicht: (2024)
Leveraging HPC Profiling & Tracing Tools to Understand the Performance of Particle-in-Cell Monte Carlo Simulations
von: Williams, Jeremy J., et al.
Veröffentlicht: (2023)
von: Williams, Jeremy J., et al.
Veröffentlicht: (2023)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
von: Ekelund, Jonah, et al.
Veröffentlicht: (2025)
von: Ekelund, Jonah, et al.
Veröffentlicht: (2025)
ARC-V: Vertical Resource Adaptivity for HPC Workloads in Containerized Environments
von: Medeiros, Daniel, et al.
Veröffentlicht: (2025)
von: Medeiros, Daniel, et al.
Veröffentlicht: (2025)
AI-coupled HPC Workflow Applications, Middleware and Performance
von: Brewer, Wes, et al.
Veröffentlicht: (2024)
von: Brewer, Wes, et al.
Veröffentlicht: (2024)
An Elastic Job Scheduler for HPC Applications on the Cloud
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
Preparing for HPC on RISC-V: Examining Vectorization and Distributed Performance of an Astrophyiscs Application with HPX and Kokkos
von: Diehl, Patrick, et al.
Veröffentlicht: (2024)
von: Diehl, Patrick, et al.
Veröffentlicht: (2024)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
Resolving Conflicts with Grace: Dynamically Concurrent Universality
von: Kuznetsov, Petr, et al.
Veröffentlicht: (2025)
von: Kuznetsov, Petr, et al.
Veröffentlicht: (2025)
Performance comparison of Dask and Apache Spark on HPC systems for Neuroimaging
von: Dugré, Mathieu, et al.
Veröffentlicht: (2024)
von: Dugré, Mathieu, et al.
Veröffentlicht: (2024)
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
von: Chazapis, Antony, et al.
Veröffentlicht: (2024)
von: Chazapis, Antony, et al.
Veröffentlicht: (2024)
Usability Evaluation of Cloud for HPC Applications
von: Sochat, Vanessa, et al.
Veröffentlicht: (2025)
von: Sochat, Vanessa, et al.
Veröffentlicht: (2025)
Host-Side Telemetry for Performance Diagnosis in Cloud and HPC GPU Infrastructure
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
A Performance Analysis of Task Scheduling for UQ Workflows on HPC Systems
von: Loi, Chung Ming, et al.
Veröffentlicht: (2025)
von: Loi, Chung Ming, et al.
Veröffentlicht: (2025)
Exploring the Viability of Unikernels for ARM-powered Edge Computing
von: Kaiser, Shahidullah, et al.
Veröffentlicht: (2024)
von: Kaiser, Shahidullah, et al.
Veröffentlicht: (2024)
Demystifying ARM SME to Optimize General Matrix Multiplications
von: Deng, Chencheng, et al.
Veröffentlicht: (2025)
von: Deng, Chencheng, et al.
Veröffentlicht: (2025)
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
von: Merzky, Andre, et al.
Veröffentlicht: (2025)
von: Merzky, Andre, et al.
Veröffentlicht: (2025)
Efficient Parameter Tuning for a Structure-Based Virtual Screening HPC Application
von: Guindani, Bruno, et al.
Veröffentlicht: (2024)
von: Guindani, Bruno, et al.
Veröffentlicht: (2024)
DOLMA: A Data Object Level Memory Disaggregation Framework for HPC Applications
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
Performance characterisation of the 64-core SG2042 RISC-V CPU for HPC
von: Brown, Nick, et al.
Veröffentlicht: (2024)
von: Brown, Nick, et al.
Veröffentlicht: (2024)
Autonomy Loops for Monitoring, Operational Data Analytics, Feedback, and Response in HPC Operations
von: Boito, Francieli, et al.
Veröffentlicht: (2024)
von: Boito, Francieli, et al.
Veröffentlicht: (2024)
IOAgent: Democratizing Trustworthy HPC I/O Performance Diagnosis Capability via LLMs
von: Egersdoerfer, Chris, et al.
Veröffentlicht: (2026)
von: Egersdoerfer, Chris, et al.
Veröffentlicht: (2026)
Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2026)
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
von: Shi, Ruimin, et al.
Veröffentlicht: (2026) -
Multi-level Memory-Centric Profiling on ARM Processors with ARM SPE
von: Miksits, Samuel, et al.
Veröffentlicht: (2024) -
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024) -
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
von: Schieffer, Gabin, et al.
Veröffentlicht: (2026) -
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2024)