Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Carrica, Vicki, Onyango, Maxwell, Alomairy, Rabab, Ringoot, Evelyne, Schloss, James, Edelman, Alan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hierarchical Recursive Precision for Accelerating Symmetric Linear Solves on MXUs
von: Carrica, Vicki, et al.
Veröffentlicht: (2026)
von: Carrica, Vicki, et al.
Veröffentlicht: (2026)
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
von: Ringoot, Evelyne, et al.
Veröffentlicht: (2025)
von: Ringoot, Evelyne, et al.
Veröffentlicht: (2025)
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
von: Ringoot, Evelyne, et al.
Veröffentlicht: (2025)
von: Ringoot, Evelyne, et al.
Veröffentlicht: (2025)
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
von: Pilliat, Emmanuel
Veröffentlicht: (2026)
von: Pilliat, Emmanuel
Veröffentlicht: (2026)
Leveraging Hardware-Aware Computation in Mixed-Precision Matrix Multiply: A Tile-Centric Approach
von: Zhang, Qiao, et al.
Veröffentlicht: (2025)
von: Zhang, Qiao, et al.
Veröffentlicht: (2025)
GPU Implementations for Midsize Integer Addition and Multiplication
von: Oancea, Cosmin E., et al.
Veröffentlicht: (2024)
von: Oancea, Cosmin E., et al.
Veröffentlicht: (2024)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
von: Wang, Hansheng, et al.
Veröffentlicht: (2025)
von: Wang, Hansheng, et al.
Veröffentlicht: (2025)
MPI Implementation Profiling for Better Application Performance
von: Shipley, Riley, et al.
Veröffentlicht: (2024)
von: Shipley, Riley, et al.
Veröffentlicht: (2024)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
von: Li, Yifan, et al.
Veröffentlicht: (2026)
von: Li, Yifan, et al.
Veröffentlicht: (2026)
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
von: Bellavita, Julian, et al.
Veröffentlicht: (2026)
von: Bellavita, Julian, et al.
Veröffentlicht: (2026)
SPES: Towards Optimizing Performance-Resource Trade-Off for Serverless Functions
von: Lee, Cheryl, et al.
Veröffentlicht: (2024)
von: Lee, Cheryl, et al.
Veröffentlicht: (2024)
Taking GPU Programming Models to Task for Performance Portability
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
Integrating Odeint Time Stepping into OpenFPM for Distributed and GPU Accelerated Numerical Solvers
von: Singh, Abhinav, et al.
Veröffentlicht: (2023)
von: Singh, Abhinav, et al.
Veröffentlicht: (2023)
Investigating Matrix Repartitioning to Address the Over- and Undersubscription Challenge for a GPU-based CFD Solver
von: Olenik, Gregor, et al.
Veröffentlicht: (2025)
von: Olenik, Gregor, et al.
Veröffentlicht: (2025)
Julia GraphBLAS with Nonblocking Execution
von: Costanza, Pascal, et al.
Veröffentlicht: (2025)
von: Costanza, Pascal, et al.
Veröffentlicht: (2025)
TorchGWAS : GPU-accelerated GWAS for thousands of quantitative phenotypes
von: Zhao, Xingzhong, et al.
Veröffentlicht: (2026)
von: Zhao, Xingzhong, et al.
Veröffentlicht: (2026)
On the energy efficiency of sparse matrix computations on multi-GPU clusters
von: Bernaschi, Massimo, et al.
Veröffentlicht: (2025)
von: Bernaschi, Massimo, et al.
Veröffentlicht: (2025)
Portability Efficiency Approach for Calculating Performance Portability
von: Marowka, Ami
Veröffentlicht: (2024)
von: Marowka, Ami
Veröffentlicht: (2024)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
Cilium and VDM -- Towards Formal Analysis of Cilium Policies
von: Kulik, Tomas, et al.
Veröffentlicht: (2024)
von: Kulik, Tomas, et al.
Veröffentlicht: (2024)
Towards an Optimized Benchmarking Platform for CI/CD Pipelines
von: Japke, Nils, et al.
Veröffentlicht: (2025)
von: Japke, Nils, et al.
Veröffentlicht: (2025)
Do Large Language Models Understand Performance Optimization?
von: Cui, Bowen, et al.
Veröffentlicht: (2025)
von: Cui, Bowen, et al.
Veröffentlicht: (2025)
Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
High-Performance Star-M SVD for Big Data Compression
von: Hussain, Md Taufique, et al.
Veröffentlicht: (2026)
von: Hussain, Md Taufique, et al.
Veröffentlicht: (2026)
Comprehensive Review of Performance Optimization Strategies for Serverless Applications on AWS Lambda
von: Bechir, Mohamed Lemine El, et al.
Veröffentlicht: (2024)
von: Bechir, Mohamed Lemine El, et al.
Veröffentlicht: (2024)
A Unifying Framework to Enable Artificial Intelligence in High Performance Computing Workflows
von: Domke, Jens, et al.
Veröffentlicht: (2025)
von: Domke, Jens, et al.
Veröffentlicht: (2025)
Cost-Performance Analysis of Cloud-Based Retail Point-of-Sale Systems: A Comparative Study of Google Cloud Platform and Microsoft Azure
von: Pagidoju, Ravi Teja
Veröffentlicht: (2026)
von: Pagidoju, Ravi Teja
Veröffentlicht: (2026)
SPUMA: a minimally invasive approach to the GPU porting of OPENFOAM
von: Bnà, Simone, et al.
Veröffentlicht: (2025)
von: Bnà, Simone, et al.
Veröffentlicht: (2025)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
von: Panova, Elena, et al.
Veröffentlicht: (2022)
von: Panova, Elena, et al.
Veröffentlicht: (2022)
GoldbachGPU: An Open Source GPU-Accelerated Framework for Verification of Goldbach's Conjecture
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026)
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026)
SGPRS: Seamless GPU Partitioning Real-Time Scheduler for Periodic Deep Learning Workloads
von: Babaei, Amir Fakhim, et al.
Veröffentlicht: (2024)
von: Babaei, Amir Fakhim, et al.
Veröffentlicht: (2024)
Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
von: Eleftherakis, Panagiotis-Eleftherios, et al.
Veröffentlicht: (2026)
von: Eleftherakis, Panagiotis-Eleftherios, et al.
Veröffentlicht: (2026)
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
von: Accordi, Gianmarco, et al.
Veröffentlicht: (2025)
von: Accordi, Gianmarco, et al.
Veröffentlicht: (2025)
Performance measurements of modern Fortran MPI applications with Score-P
von: Corbin, Gregor
Veröffentlicht: (2025)
von: Corbin, Gregor
Veröffentlicht: (2025)
BandPilot: Towards Performance- and Contention-Aware GPU Dispatching in AI Clusters
von: Zhang, Kunming, et al.
Veröffentlicht: (2025)
von: Zhang, Kunming, et al.
Veröffentlicht: (2025)
LAPIS: A Performance Portable, High Productivity Compiler Framework
von: Kelley, Brian, et al.
Veröffentlicht: (2025)
von: Kelley, Brian, et al.
Veröffentlicht: (2025)
HPDR: High-Performance Portable Scientific Data Reduction Framework
von: Chen, Jieyang, et al.
Veröffentlicht: (2025)
von: Chen, Jieyang, et al.
Veröffentlicht: (2025)
XaaS Containers: Performance-Portable Representation With Source and IR Containers
von: Copik, Marcin, et al.
Veröffentlicht: (2025)
von: Copik, Marcin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hierarchical Recursive Precision for Accelerating Symmetric Linear Solves on MXUs
von: Carrica, Vicki, et al.
Veröffentlicht: (2026) -
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
von: Ringoot, Evelyne, et al.
Veröffentlicht: (2025) -
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
von: Ringoot, Evelyne, et al.
Veröffentlicht: (2025) -
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025) -
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
von: Pilliat, Emmanuel
Veröffentlicht: (2026)