GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shan, Baodi, Araya-Polo, Mauricio, Chapman, Barbara |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
DiOMP-Offloading: Toward Portable Distributed Heterogeneous OpenMP
von: Shan, Baodi, et al.
Veröffentlicht: (2025)
von: Shan, Baodi, et al.
Veröffentlicht: (2025)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
von: Merzky, Andre, et al.
Veröffentlicht: (2025)
von: Merzky, Andre, et al.
Veröffentlicht: (2025)
Towards an Adaptive Runtime System for Cloud-Native HPC
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
A Portable Framework for Accelerating Stencil Computations on Modern Node Architectures
von: Sai, Ryuichi, et al.
Veröffentlicht: (2023)
von: Sai, Ryuichi, et al.
Veröffentlicht: (2023)
Host-Side Telemetry for Performance Diagnosis in Cloud and HPC GPU Infrastructure
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
Characterizing the Impact of Congestion in Modern HPC Interconnects
von: Piarulli, Lorenzo, et al.
Veröffentlicht: (2026)
von: Piarulli, Lorenzo, et al.
Veröffentlicht: (2026)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
von: Xu, Zhihao, et al.
Veröffentlicht: (2025)
von: Xu, Zhihao, et al.
Veröffentlicht: (2025)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
von: Jain, Rutwik, et al.
Veröffentlicht: (2026)
von: Jain, Rutwik, et al.
Veröffentlicht: (2026)
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
von: Chazapis, Antony, et al.
Veröffentlicht: (2024)
von: Chazapis, Antony, et al.
Veröffentlicht: (2024)
A Performance Analysis of Task Scheduling for UQ Workflows on HPC Systems
von: Loi, Chung Ming, et al.
Veröffentlicht: (2025)
von: Loi, Chung Ming, et al.
Veröffentlicht: (2025)
MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems
von: Kwon, Miryeong, et al.
Veröffentlicht: (2025)
von: Kwon, Miryeong, et al.
Veröffentlicht: (2025)
Parallel Paradigms in Modern HPC: A Comparative Analysis of MPI, OpenMP, and CUDA
von: ALHafez, Nizar, et al.
Veröffentlicht: (2025)
von: ALHafez, Nizar, et al.
Veröffentlicht: (2025)
FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
von: Park, Seongyeon, et al.
Veröffentlicht: (2025)
von: Park, Seongyeon, et al.
Veröffentlicht: (2025)
Optimizing Bloom Filters for Modern GPU Architectures
von: Jünger, Daniel, et al.
Veröffentlicht: (2025)
von: Jünger, Daniel, et al.
Veröffentlicht: (2025)
AI-coupled HPC Workflow Applications, Middleware and Performance
von: Brewer, Wes, et al.
Veröffentlicht: (2024)
von: Brewer, Wes, et al.
Veröffentlicht: (2024)
GTaP: A GPU-Resident Fork-Join Task-Parallel Runtime with a Pragma-Based Interface
von: Maeda, Yuki, et al.
Veröffentlicht: (2026)
von: Maeda, Yuki, et al.
Veröffentlicht: (2026)
On the Convergence of Malleability and the HPC PowerStack: Exploiting Dynamism in Over-Provisioned and Power-Constrained HPC Systems
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
SIREN: Software Identification and Recognition in HPC Systems
von: Jakobsche, Thomas, et al.
Veröffentlicht: (2025)
von: Jakobsche, Thomas, et al.
Veröffentlicht: (2025)
Towards Experiment Execution in Support of Community Benchmark Workflows for HPC
von: von Laszewski, Gregor, et al.
Veröffentlicht: (2025)
von: von Laszewski, Gregor, et al.
Veröffentlicht: (2025)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
Performance comparison of Dask and Apache Spark on HPC systems for Neuroimaging
von: Dugré, Mathieu, et al.
Veröffentlicht: (2024)
von: Dugré, Mathieu, et al.
Veröffentlicht: (2024)
Computational Performance and Energy Efficiency of ARM based HPC servers
von: Schirmer, Oskar
Veröffentlicht: (2024)
von: Schirmer, Oskar
Veröffentlicht: (2024)
Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications
von: Merzky, Andre, et al.
Veröffentlicht: (2025)
von: Merzky, Andre, et al.
Veröffentlicht: (2025)
A New Execution Model and Executor for Adaptively Optimizing the Performance of Parallel Algorithms Using HPX Runtime System
von: Mohammadiporshokooh, Karame, et al.
Veröffentlicht: (2025)
von: Mohammadiporshokooh, Karame, et al.
Veröffentlicht: (2025)
Automatic Tracing in Task-Based Runtime Systems
von: Yadav, Rohan, et al.
Veröffentlicht: (2024)
von: Yadav, Rohan, et al.
Veröffentlicht: (2024)
TurboFFT: A High-Performance Fast Fourier Transform with Fault Tolerance on GPU
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
Report on Challenges of Practical Reproducibility for Systems and HPC Computer Science
von: Keahey, Kate, et al.
Veröffentlicht: (2025)
von: Keahey, Kate, et al.
Veröffentlicht: (2025)
Optimizing Allreduce Operations for Modern Heterogeneous Architectures with Multiple Processes per GPU
von: Adams, Michael, et al.
Veröffentlicht: (2025)
von: Adams, Michael, et al.
Veröffentlicht: (2025)
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
von: Shi, Ruimin, et al.
Veröffentlicht: (2025)
von: Shi, Ruimin, et al.
Veröffentlicht: (2025)
Performance characterisation of the 64-core SG2042 RISC-V CPU for HPC
von: Brown, Nick, et al.
Veröffentlicht: (2024)
von: Brown, Nick, et al.
Veröffentlicht: (2024)
Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems
von: Knorr, Fabian, et al.
Veröffentlicht: (2025)
von: Knorr, Fabian, et al.
Veröffentlicht: (2025)
Privacy-Preserving Sharing of Data Analytics Runtime Metrics for Performance Modeling
von: Will, Jonathan, et al.
Veröffentlicht: (2024)
von: Will, Jonathan, et al.
Veröffentlicht: (2024)
A Parallel and Highly-Portable HPC Poisson Solver: Preconditioned Bi-CGSTAB with alpaka
von: Pennati, Luca, et al.
Veröffentlicht: (2025)
von: Pennati, Luca, et al.
Veröffentlicht: (2025)
High-Performance N-Queens Solver on GPU: Iterative DFS with Zero Bank Conflicts
von: Yao, Guangchao, et al.
Veröffentlicht: (2025)
von: Yao, Guangchao, et al.
Veröffentlicht: (2025)
Beyond Pre-Training: The Full Lifecycle of Foundation Models on HPC Systems
von: Conciatore, Dino, et al.
Veröffentlicht: (2026)
von: Conciatore, Dino, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures
von: Shan, Baodi, et al.
Veröffentlicht: (2024) -
DiOMP-Offloading: Toward Portable Distributed Heterogeneous OpenMP
von: Shan, Baodi, et al.
Veröffentlicht: (2025) -
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
von: Shan, Baodi, et al.
Veröffentlicht: (2024) -
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
von: Merzky, Andre, et al.
Veröffentlicht: (2025) -
Towards an Adaptive Runtime System for Cloud-Native HPC
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)