MPI Implementation Profiling for Better Application Performance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shipley, Riley, Hooten, Garrett, Boehme, David, Schafer, Derek, Skjellum, Anthony, Pearce, Olga |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Caliper and Benchpark to Analyze MPI Communication Patterns: Insights from AMG2023, Kripke, and Laghos
von: Nansamba, Grace, et al.
Veröffentlicht: (2025)
von: Nansamba, Grace, et al.
Veröffentlicht: (2025)
LibProf: A Python Profiler for Improving Cold Start Performance in Serverless Applications
von: Tariq, Syed Salauddin Mohammad, et al.
Veröffentlicht: (2024)
von: Tariq, Syed Salauddin Mohammad, et al.
Veröffentlicht: (2024)
When Should I Run My Application Benchmark?: Studying Cloud Performance Variability for the Case of Stream Processing Applications
von: Henning, Sören, et al.
Veröffentlicht: (2025)
von: Henning, Sören, et al.
Veröffentlicht: (2025)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
Performance measurements of modern Fortran MPI applications with Score-P
von: Corbin, Gregor
Veröffentlicht: (2025)
von: Corbin, Gregor
Veröffentlicht: (2025)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
von: Li, Junjie
Veröffentlicht: (2024)
von: Li, Junjie
Veröffentlicht: (2024)
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
von: Wang, Zirui, et al.
Veröffentlicht: (2026)
von: Wang, Zirui, et al.
Veröffentlicht: (2026)
High-level Stream Processing: A Complementary Analysis of Fault Recovery
von: Vogel, Adriano, et al.
Veröffentlicht: (2024)
von: Vogel, Adriano, et al.
Veröffentlicht: (2024)
Optimizing OpenFaaS on Kubernetes: Comparative Analysis of Language Runtimes and Cluster Distributions
von: Ataie, Ehsan, et al.
Veröffentlicht: (2026)
von: Ataie, Ehsan, et al.
Veröffentlicht: (2026)
Where Should I Deploy My Contracts? A Practical Experience Report
von: Lazăr, Cătălina, et al.
Veröffentlicht: (2025)
von: Lazăr, Cătălina, et al.
Veröffentlicht: (2025)
Should I Run My Cloud Benchmark on Black Friday?
von: Henning, Sören, et al.
Veröffentlicht: (2025)
von: Henning, Sören, et al.
Veröffentlicht: (2025)
Automated MPI-X code generation for scalable finite-difference solvers
von: Bisbas, George, et al.
Veröffentlicht: (2023)
von: Bisbas, George, et al.
Veröffentlicht: (2023)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
von: Tuteja, Keshvi, et al.
Veröffentlicht: (2025)
von: Tuteja, Keshvi, et al.
Veröffentlicht: (2025)
Rethinking Performance Analysis for Configurable Software Systems: A Case Study from a Fitness Landscape Perspective
von: Huang, Mingyu, et al.
Veröffentlicht: (2024)
von: Huang, Mingyu, et al.
Veröffentlicht: (2024)
Synthesizing Proxy Applications for MPI Programs
von: Luo, Jiyu, et al.
Veröffentlicht: (2023)
von: Luo, Jiyu, et al.
Veröffentlicht: (2023)
Understanding GPU Triggering APIs for MPI+X Communication
von: Bridges, Patrick G., et al.
Veröffentlicht: (2024)
von: Bridges, Patrick G., et al.
Veröffentlicht: (2024)
Leveraging AI for Productive and Trustworthy HPC Software: Challenges and Research Directions
von: Teranishi, Keita, et al.
Veröffentlicht: (2025)
von: Teranishi, Keita, et al.
Veröffentlicht: (2025)
MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
von: Wen, Zhongzhen, et al.
Veröffentlicht: (2025)
von: Wen, Zhongzhen, et al.
Veröffentlicht: (2025)
A Delta-Aware Orchestration Framework for Scalable Multi-Agent Edge Computing
von: Singh, Samaresh Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Samaresh Kumar, et al.
Veröffentlicht: (2026)
Optimized thread-block arrangement in a GPU implementation of a linear solver for atmospheric chemistry mechanisms
von: Ruiz, Christian Guzman, et al.
Veröffentlicht: (2024)
von: Ruiz, Christian Guzman, et al.
Veröffentlicht: (2024)
Evaluating Asynchronous Semantics in Trace-Discovered Resilience Models: A Case Study on the OpenTelemetry Demo
von: Krasnovsky, Anatoly A.
Veröffentlicht: (2025)
von: Krasnovsky, Anatoly A.
Veröffentlicht: (2025)
Emergence-as-Code for Self-Governing Reliable Systems
von: Krasnovsky, Anatoly A.
Veröffentlicht: (2026)
von: Krasnovsky, Anatoly A.
Veröffentlicht: (2026)
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
von: Wen, Zhongzhen, et al.
Veröffentlicht: (2026)
von: Wen, Zhongzhen, et al.
Veröffentlicht: (2026)
Evaluating Fault Tolerance and Scalability in Distributed File Systems: A Case Study of GFS, HDFS, and MinIO
von: Malhotra, Shubham, et al.
Veröffentlicht: (2025)
von: Malhotra, Shubham, et al.
Veröffentlicht: (2025)
Adaptive Protein Design Protocols and Middleware
von: Alsaadi, Aymen, et al.
Veröffentlicht: (2025)
von: Alsaadi, Aymen, et al.
Veröffentlicht: (2025)
MPI Errors Detection using GNN Embedding and Vector Embedding over LLVM IR
von: Karchi, Jad El, et al.
Veröffentlicht: (2024)
von: Karchi, Jad El, et al.
Veröffentlicht: (2024)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
von: Panova, Elena, et al.
Veröffentlicht: (2022)
von: Panova, Elena, et al.
Veröffentlicht: (2022)
LLM-HPC++: Evaluating LLM-Generated Modern C++ and MPI+OpenMP Codes for Scalable Mandelbrot Set Computation
von: Diehl, Patrick, et al.
Veröffentlicht: (2025)
von: Diehl, Patrick, et al.
Veröffentlicht: (2025)
CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe
von: Saba, Tara, et al.
Veröffentlicht: (2026)
von: Saba, Tara, et al.
Veröffentlicht: (2026)
The Case for ABI Interoperability in a Fault Tolerant MPI
von: Xu, Yao, et al.
Veröffentlicht: (2025)
von: Xu, Yao, et al.
Veröffentlicht: (2025)
Easy Acceleration with Distributed Arrays
von: Kepner, Jeremy, et al.
Veröffentlicht: (2025)
von: Kepner, Jeremy, et al.
Veröffentlicht: (2025)
NApy: Efficient Statistics in Python for Large-Scale Heterogeneous Data with Enhanced Support for Missing Data
von: Woller, Fabian, et al.
Veröffentlicht: (2025)
von: Woller, Fabian, et al.
Veröffentlicht: (2025)
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
von: Bridges, Patrick G., et al.
Veröffentlicht: (2026)
von: Bridges, Patrick G., et al.
Veröffentlicht: (2026)
Comprehensive Review of Performance Optimization Strategies for Serverless Applications on AWS Lambda
von: Bechir, Mohamed Lemine El, et al.
Veröffentlicht: (2024)
von: Bechir, Mohamed Lemine El, et al.
Veröffentlicht: (2024)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
On the energy efficiency of sparse matrix computations on multi-GPU clusters
von: Bernaschi, Massimo, et al.
Veröffentlicht: (2025)
von: Bernaschi, Massimo, et al.
Veröffentlicht: (2025)
A Communication Avoiding and Reducing Algorithm for Symmetric Eigenproblem for Very Small Matrices
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
von: Bergach, Mohamed Amine
Veröffentlicht: (2026)
von: Bergach, Mohamed Amine
Veröffentlicht: (2026)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
Umbilical Choir: Automated Live Testing for Edge-To-Cloud FaaS Applications
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2025)
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leveraging Caliper and Benchpark to Analyze MPI Communication Patterns: Insights from AMG2023, Kripke, and Laghos
von: Nansamba, Grace, et al.
Veröffentlicht: (2025) -
LibProf: A Python Profiler for Improving Cold Start Performance in Serverless Applications
von: Tariq, Syed Salauddin Mohammad, et al.
Veröffentlicht: (2024) -
When Should I Run My Application Benchmark?: Studying Cloud Performance Variability for the Case of Stream Processing Applications
von: Henning, Sören, et al.
Veröffentlicht: (2025) -
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
von: Nichols, Daniel, et al.
Veröffentlicht: (2025) -
Performance measurements of modern Fortran MPI applications with Score-P
von: Corbin, Gregor
Veröffentlicht: (2025)