A Portable Framework for Accelerating Stencil Computations on Modern Node Architectures
Fuente:
arXiv
Saved in:
| Main Authors: | Sai, Ryuichi, Mellor-Crummey, John, Xu, Jinfan, Araya-Polo, Mauricio |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures
by: Shan, Baodi, et al.
Published: (2024)
by: Shan, Baodi, et al.
Published: (2024)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
by: Xia, Yuning, et al.
Published: (2026)
by: Xia, Yuning, et al.
Published: (2026)
DiOMP-Offloading: Toward Portable Distributed Heterogeneous OpenMP
by: Shan, Baodi, et al.
Published: (2025)
by: Shan, Baodi, et al.
Published: (2025)
GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems
by: Shan, Baodi, et al.
Published: (2026)
by: Shan, Baodi, et al.
Published: (2026)
Stencil Matrixization
by: Zhao, Wenxuan, et al.
Published: (2023)
by: Zhao, Wenxuan, et al.
Published: (2023)
Do We Need Tensor Cores for Stencil Computations?
by: Gu, Qiqi, et al.
Published: (2026)
by: Gu, Qiqi, et al.
Published: (2026)
Stencil Computations on Tenstorrent Wormhole
by: Piarulli, Lorenzo, et al.
Published: (2026)
by: Piarulli, Lorenzo, et al.
Published: (2026)
SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping
by: GU, Qiqi, et al.
Published: (2025)
by: GU, Qiqi, et al.
Published: (2025)
Persistent and Partitioned MPI for Stencil Communication
by: Collom, Gerald, et al.
Published: (2025)
by: Collom, Gerald, et al.
Published: (2025)
An Adaptive Distributed Stencil Abstraction for GPUs
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
by: Shan, Baodi, et al.
Published: (2024)
by: Shan, Baodi, et al.
Published: (2024)
To Repair or Not to Repair: Assessing Fault Resilience in MPI Stencil Applications
by: Rocco, Roberto, et al.
Published: (2024)
by: Rocco, Roberto, et al.
Published: (2024)
Akita: A High Usability Simulation Framework for Computer Architecture
by: Jannat, Sabila Al, et al.
Published: (2026)
by: Jannat, Sabila Al, et al.
Published: (2026)
MMStencil: Optimizing High-order Stencils on Multicore CPU using Matrix Unit
by: Wang, Yinuo, et al.
Published: (2025)
by: Wang, Yinuo, et al.
Published: (2025)
LAPIS: A Performance Portable, High Productivity Compiler Framework
by: Kelley, Brian, et al.
Published: (2025)
by: Kelley, Brian, et al.
Published: (2025)
Portability Efficiency Approach for Calculating Performance Portability
by: Marowka, Ami
Published: (2024)
by: Marowka, Ami
Published: (2024)
QPU Micro-Kernels for Stencil Computation
by: Markidis, Stefano, et al.
Published: (2025)
by: Markidis, Stefano, et al.
Published: (2025)
HPDR: High-Performance Portable Scientific Data Reduction Framework
by: Chen, Jieyang, et al.
Published: (2025)
by: Chen, Jieyang, et al.
Published: (2025)
Optimizing Bloom Filters for Modern GPU Architectures
by: Jünger, Daniel, et al.
Published: (2025)
by: Jünger, Daniel, et al.
Published: (2025)
Modern Computing: Vision and Challenges
by: Gill, Sukhpal Singh, et al.
Published: (2024)
by: Gill, Sukhpal Singh, et al.
Published: (2024)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
by: Rose, Martin, et al.
Published: (2025)
by: Rose, Martin, et al.
Published: (2025)
Portable, heterogeneous ensemble workflows at scale using libEnsemble
by: Hudson, Stephen, et al.
Published: (2024)
by: Hudson, Stephen, et al.
Published: (2024)
CHIRON: Accelerating Node Synchronization without Security Trade-offs in Distributed Ledgers
by: Neiheiser, Ray, et al.
Published: (2024)
by: Neiheiser, Ray, et al.
Published: (2024)
Automating Multi-Tenancy Performance Evaluation on Edge Compute Nodes
by: Georgiou, Joanna, et al.
Published: (2025)
by: Georgiou, Joanna, et al.
Published: (2025)
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
by: Villalobos, Johansell, et al.
Published: (2025)
by: Villalobos, Johansell, et al.
Published: (2025)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
by: Sojoodi, Amirhossein, et al.
Published: (2026)
by: Sojoodi, Amirhossein, et al.
Published: (2026)
Syndeo: Portable Ray Clusters with Secure Containerization
by: Li, William, et al.
Published: (2024)
by: Li, William, et al.
Published: (2024)
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
by: Tramm, John, et al.
Published: (2024)
by: Tramm, John, et al.
Published: (2024)
Characterizing the Performance of Accelerated Jetson Edge Devices for Training Deep Learning Models
by: K., Prashanthi S., et al.
Published: (2025)
by: K., Prashanthi S., et al.
Published: (2025)
Optimizing Allreduce Operations for Modern Heterogeneous Architectures with Multiple Processes per GPU
by: Adams, Michael, et al.
Published: (2025)
by: Adams, Michael, et al.
Published: (2025)
Comparing the Run-time Behavior of Modern PDES Engines on Alternative Hardware Architectures
by: Marotta, Romolo, et al.
Published: (2025)
by: Marotta, Romolo, et al.
Published: (2025)
BlockRaFT: A Distributed Framework for Fault-Tolerant and Scalable Blockchain Nodes
by: Piduguralla, Manaswini, et al.
Published: (2026)
by: Piduguralla, Manaswini, et al.
Published: (2026)
Advancing Blockchain Scalability: A Linear Optimization Framework for Diversified Node Allocation in Shards
by: Assmann, Björn, et al.
Published: (2024)
by: Assmann, Björn, et al.
Published: (2024)
Layout-Agnostic MPI Abstraction for Distributed Computing in Modern C++
by: Klepl, Jiří, et al.
Published: (2025)
by: Klepl, Jiří, et al.
Published: (2025)
PSI/J: A Portable Interface for Submitting, Monitoring, and Managing Jobs
by: Hategan-Marandiuc, Mihael, et al.
Published: (2023)
by: Hategan-Marandiuc, Mihael, et al.
Published: (2023)
miniLB: A Performance Portability Study of Lattice-Boltzmann Simulations
by: Crisci, Luigi, et al.
Published: (2024)
by: Crisci, Luigi, et al.
Published: (2024)
Serverless Computing: Architecture, Concepts, and Applications
by: Ghorbian, Mohsen, et al.
Published: (2025)
by: Ghorbian, Mohsen, et al.
Published: (2025)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
by: Villalobos, Johansell, et al.
Published: (2025)
by: Villalobos, Johansell, et al.
Published: (2025)
Stencil Computations on Cerebras Wafer-Scale Engine
by: Belli, Elia, et al.
Published: (2026)
by: Belli, Elia, et al.
Published: (2026)
Maple: A Multi-agent System for Portable Deep Learning across Clusters
by: Wu, Molang, et al.
Published: (2025)
by: Wu, Molang, et al.
Published: (2025)
Similar Items
-
Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures
by: Shan, Baodi, et al.
Published: (2024) -
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
by: Xia, Yuning, et al.
Published: (2026) -
DiOMP-Offloading: Toward Portable Distributed Heterogeneous OpenMP
by: Shan, Baodi, et al.
Published: (2025) -
GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems
by: Shan, Baodi, et al.
Published: (2026) -
Stencil Matrixization
by: Zhao, Wenxuan, et al.
Published: (2023)