Automated MPI-X code generation for scalable finite-difference solvers
Fuente:
arXiv
Saved in:
| Main Authors: | Bisbas, George, Nelson, Rhodri, Louboutin, Mathias, Luporini, Fabio, Kelly, Paul H. J., Gorman, Gerard |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MPI Implementation Profiling for Better Application Performance
by: Shipley, Riley, et al.
Published: (2024)
by: Shipley, Riley, et al.
Published: (2024)
Performance measurements of modern Fortran MPI applications with Score-P
by: Corbin, Gregor
Published: (2025)
by: Corbin, Gregor
Published: (2025)
Synthesizing Proxy Applications for MPI Programs
by: Luo, Jiyu, et al.
Published: (2023)
by: Luo, Jiyu, et al.
Published: (2023)
Optimized thread-block arrangement in a GPU implementation of a linear solver for atmospheric chemistry mechanisms
by: Ruiz, Christian Guzman, et al.
Published: (2024)
by: Ruiz, Christian Guzman, et al.
Published: (2024)
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
by: Wang, Zirui, et al.
Published: (2026)
by: Wang, Zirui, et al.
Published: (2026)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
High-level Stream Processing: A Complementary Analysis of Fault Recovery
by: Vogel, Adriano, et al.
Published: (2024)
by: Vogel, Adriano, et al.
Published: (2024)
LibProf: A Python Profiler for Improving Cold Start Performance in Serverless Applications
by: Tariq, Syed Salauddin Mohammad, et al.
Published: (2024)
by: Tariq, Syed Salauddin Mohammad, et al.
Published: (2024)
Optimizing OpenFaaS on Kubernetes: Comparative Analysis of Language Runtimes and Cluster Distributions
by: Ataie, Ehsan, et al.
Published: (2026)
by: Ataie, Ehsan, et al.
Published: (2026)
On the energy efficiency of sparse matrix computations on multi-GPU clusters
by: Bernaschi, Massimo, et al.
Published: (2025)
by: Bernaschi, Massimo, et al.
Published: (2025)
When Should I Run My Application Benchmark?: Studying Cloud Performance Variability for the Case of Stream Processing Applications
by: Henning, Sören, et al.
Published: (2025)
by: Henning, Sören, et al.
Published: (2025)
Where Should I Deploy My Contracts? A Practical Experience Report
by: Lazăr, Cătălina, et al.
Published: (2025)
by: Lazăr, Cătălina, et al.
Published: (2025)
NApy: Efficient Statistics in Python for Large-Scale Heterogeneous Data with Enhanced Support for Missing Data
by: Woller, Fabian, et al.
Published: (2025)
by: Woller, Fabian, et al.
Published: (2025)
A Communication Avoiding and Reducing Algorithm for Symmetric Eigenproblem for Very Small Matrices
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
by: Nichols, Daniel, et al.
Published: (2025)
by: Nichols, Daniel, et al.
Published: (2025)
Should I Run My Cloud Benchmark on Black Friday?
by: Henning, Sören, et al.
Published: (2025)
by: Henning, Sören, et al.
Published: (2025)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
by: Panova, Elena, et al.
Published: (2022)
by: Panova, Elena, et al.
Published: (2022)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
by: Tu, Jiqun, et al.
Published: (2026)
by: Tu, Jiqun, et al.
Published: (2026)
Improving the scalability of a high-order atmospheric dynamics solver based on the deal.II library
by: Orlando, Giuseppe, et al.
Published: (2025)
by: Orlando, Giuseppe, et al.
Published: (2025)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
by: Tuteja, Keshvi, et al.
Published: (2025)
by: Tuteja, Keshvi, et al.
Published: (2025)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
by: Li, Junjie
Published: (2024)
by: Li, Junjie
Published: (2024)
Performance and scaling of the LFRic weather and climate model on different generations of HPE Cray EX supercomputers
by: Bull, J. Mark, et al.
Published: (2024)
by: Bull, J. Mark, et al.
Published: (2024)
Automated Programmatic Performance Analysis of Parallel Programs
by: Cankur, Onur, et al.
Published: (2024)
by: Cankur, Onur, et al.
Published: (2024)
MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
by: Wen, Zhongzhen, et al.
Published: (2025)
by: Wen, Zhongzhen, et al.
Published: (2025)
A Delta-Aware Orchestration Framework for Scalable Multi-Agent Edge Computing
by: Singh, Samaresh Kumar, et al.
Published: (2026)
by: Singh, Samaresh Kumar, et al.
Published: (2026)
Rethinking Performance Analysis for Configurable Software Systems: A Case Study from a Fitness Landscape Perspective
by: Huang, Mingyu, et al.
Published: (2024)
by: Huang, Mingyu, et al.
Published: (2024)
Evaluating Asynchronous Semantics in Trace-Discovered Resilience Models: A Case Study on the OpenTelemetry Demo
by: Krasnovsky, Anatoly A.
Published: (2025)
by: Krasnovsky, Anatoly A.
Published: (2025)
GoldbachGPU: An Open Source GPU-Accelerated Framework for Verification of Goldbach's Conjecture
by: Llorente-Saguer, Isaac
Published: (2026)
by: Llorente-Saguer, Isaac
Published: (2026)
Emergence-as-Code for Self-Governing Reliable Systems
by: Krasnovsky, Anatoly A.
Published: (2026)
by: Krasnovsky, Anatoly A.
Published: (2026)
Hierarchical Recursive Precision for Accelerating Symmetric Linear Solves on MXUs
by: Carrica, Vicki, et al.
Published: (2026)
by: Carrica, Vicki, et al.
Published: (2026)
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
by: Wen, Zhongzhen, et al.
Published: (2026)
by: Wen, Zhongzhen, et al.
Published: (2026)
Evaluating Fault Tolerance and Scalability in Distributed File Systems: A Case Study of GFS, HDFS, and MinIO
by: Malhotra, Shubham, et al.
Published: (2025)
by: Malhotra, Shubham, et al.
Published: (2025)
Leveraging AI for Productive and Trustworthy HPC Software: Challenges and Research Directions
by: Teranishi, Keita, et al.
Published: (2025)
by: Teranishi, Keita, et al.
Published: (2025)
Adaptive Protein Design Protocols and Middleware
by: Alsaadi, Aymen, et al.
Published: (2025)
by: Alsaadi, Aymen, et al.
Published: (2025)
Enabling MPI communication within Numba/LLVM JIT-compiled Python code using numba-mpi v1.0
by: Derlatka, Kacper, et al.
Published: (2024)
by: Derlatka, Kacper, et al.
Published: (2024)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
by: McDonald, Jesse, et al.
Published: (2024)
by: McDonald, Jesse, et al.
Published: (2024)
FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents
by: Battaglini-Fischer, Sándor, et al.
Published: (2025)
by: Battaglini-Fischer, Sándor, et al.
Published: (2025)
Ridgeline: A 2D Roofline Model for Distributed Systems
by: Checconi, Fabio, et al.
Published: (2022)
by: Checconi, Fabio, et al.
Published: (2022)
Accelerating Particle-in-Cell Monte Carlo Simulations with MPI, OpenMP/OpenACC and Asynchronous Multi-GPU Programming
by: Williams, Jeremy J., et al.
Published: (2024)
by: Williams, Jeremy J., et al.
Published: (2024)
Similar Items
-
MPI Implementation Profiling for Better Application Performance
by: Shipley, Riley, et al.
Published: (2024) -
Performance measurements of modern Fortran MPI applications with Score-P
by: Corbin, Gregor
Published: (2025) -
Synthesizing Proxy Applications for MPI Programs
by: Luo, Jiyu, et al.
Published: (2023) -
Optimized thread-block arrangement in a GPU implementation of a linear solver for atmospheric chemistry mechanisms
by: Ruiz, Christian Guzman, et al.
Published: (2024) -
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
by: Wang, Zirui, et al.
Published: (2026)