Effective implementation of the High Performance Conjugate Gradient benchmark on GraphBLAS
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Scolari, Alberto, Yzelman, Albert-Jan |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Nonlinear spectral clustering with C++ GraphBLAS
par: Pasadakis, Dimosthenis, et autres
Publié: (2026)
par: Pasadakis, Dimosthenis, et autres
Publié: (2026)
Hypersparse Traffic Matrices from Suricata Network Flows using GraphBLAS
par: Houle, Michael, et autres
Publié: (2024)
par: Houle, Michael, et autres
Publié: (2024)
Julia GraphBLAS with Nonblocking Execution
par: Costanza, Pascal, et autres
Publié: (2025)
par: Costanza, Pascal, et autres
Publié: (2025)
Faster Distributed Inference-Only Recommender Systems via Bounded Lag Synchronous Collectives
par: Dichev, Kiril, et autres
Publié: (2025)
par: Dichev, Kiril, et autres
Publié: (2025)
HiCR, an Abstract Model for Distributed Heterogeneous Programming
par: Martin, Sergio Miguel, et autres
Publié: (2025)
par: Martin, Sergio Miguel, et autres
Publié: (2025)
Performance optimization of BLAS algorithms with band matrices for RISC-V processors
par: Pirova, Anna, et autres
Publié: (2025)
par: Pirova, Anna, et autres
Publié: (2025)
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
par: Swann, Ryan, et autres
Publié: (2025)
par: Swann, Ryan, et autres
Publié: (2025)
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
par: Li, Junjie, et autres
Publié: (2024)
par: Li, Junjie, et autres
Publié: (2024)
Mapping Parallel Matrix Multiplication in GotoBLAS2 to the AMD Versal ACAP for Deep Learning
par: Lei, Jie, et autres
Publié: (2024)
par: Lei, Jie, et autres
Publié: (2024)
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
par: Thüring, Tim, et autres
Publié: (2026)
par: Thüring, Tim, et autres
Publié: (2026)
A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
par: Liu, Hang, et autres
Publié: (2025)
par: Liu, Hang, et autres
Publié: (2025)
Developing a BLAS library for the AMD AI Engine
par: Laan, Tristan, et autres
Publié: (2024)
par: Laan, Tristan, et autres
Publié: (2024)
Continuous benchmarking: Keeping pace with an evolving ecosystem of models and technologies
par: Vogelsang, Jan, et autres
Publié: (2026)
par: Vogelsang, Jan, et autres
Publié: (2026)
Testing and benchmarking emerging supercomputers via the MFC flow solver
par: Wilfong, Benjamin, et autres
Publié: (2025)
par: Wilfong, Benjamin, et autres
Publié: (2025)
MPCGPU: Real-Time Nonlinear Model Predictive Control through Preconditioned Conjugate Gradient on the GPU
par: Adabag, Emre, et autres
Publié: (2023)
par: Adabag, Emre, et autres
Publié: (2023)
Performance Comparison of Graph Representations Which Support Dynamic Graph Updates
par: Sahu, Subhajit
Publié: (2025)
par: Sahu, Subhajit
Publié: (2025)
Machine-Learning-Driven Runtime Optimization of BLAS Level 3 on Modern Multi-Core Systems
par: Xia, Yufan, et autres
Publié: (2024)
par: Xia, Yufan, et autres
Publié: (2024)
Exploration on Highly Dynamic Graphs
par: Saxena, Ashish, et autres
Publié: (2026)
par: Saxena, Ashish, et autres
Publié: (2026)
Task queue implementation for edge computing platform
par: Maksimovic, Veljko, et autres
Publié: (2024)
par: Maksimovic, Veljko, et autres
Publié: (2024)
Evaluation of computational and energy performance in matrix multiplication algorithms on CPU and GPU using MKL, cuBLAS and SYCL
par: Torres, L. A., et autres
Publié: (2024)
par: Torres, L. A., et autres
Publié: (2024)
FastGraph: Optimized GPU-Enabled Algorithms for Fast Graph Building and Message Passing
par: Agarwal, Aarush, et autres
Publié: (2025)
par: Agarwal, Aarush, et autres
Publié: (2025)
Use Cases for High Performance Research Desktops
par: Henschel, Robert, et autres
Publié: (2024)
par: Henschel, Robert, et autres
Publié: (2024)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
par: Ekelund, Jonah, et autres
Publié: (2025)
par: Ekelund, Jonah, et autres
Publié: (2025)
The Fused Kernel Library: A C++ API to Develop Highly-Efficient GPU Libraries
par: Amoros, Oscar, et autres
Publié: (2025)
par: Amoros, Oscar, et autres
Publié: (2025)
Towards observability of scientific applications
par: Balis, Bartosz, et autres
Publié: (2024)
par: Balis, Bartosz, et autres
Publié: (2024)
Cooperative Gradient Coding
par: Weng, Shudi, et autres
Publié: (2025)
par: Weng, Shudi, et autres
Publié: (2025)
Automatic Metadata Capture and Processing for High-Performance Workflows
par: Shpilker, Polina, et autres
Publié: (2025)
par: Shpilker, Polina, et autres
Publié: (2025)
Raptr: Prefix Consensus for Robust High-Performance BFT
par: Tonkikh, Andrei, et autres
Publié: (2025)
par: Tonkikh, Andrei, et autres
Publié: (2025)
HP2C-DT: High-Precision High-Performance Computer-enabled Digital Twin
par: Iraola, E., et autres
Publié: (2025)
par: Iraola, E., et autres
Publié: (2025)
Robustness and Accuracy in Pipelined Bi-Conjugate Gradient Stabilized Method: A Comparative Study
par: Havdiak, Mykhailo, et autres
Publié: (2024)
par: Havdiak, Mykhailo, et autres
Publié: (2024)
LAPIS: A Performance Portable, High Productivity Compiler Framework
par: Kelley, Brian, et autres
Publié: (2025)
par: Kelley, Brian, et autres
Publié: (2025)
High-Performance Parallelization of Dijkstra's Algorithm Using MPI and CUDA
par: Song, Boyang
Publié: (2025)
par: Song, Boyang
Publié: (2025)
HPDR: High-Performance Portable Scientific Data Reduction Framework
par: Chen, Jieyang, et autres
Publié: (2025)
par: Chen, Jieyang, et autres
Publié: (2025)
Modular Architecture for High-Performance and Low Overhead Data Transfers
par: Swargo, Rasman Mubtasim, et autres
Publié: (2025)
par: Swargo, Rasman Mubtasim, et autres
Publié: (2025)
Cppless: Single-Source and High-Performance Serverless Programming in C++
par: Copik, Marcin, et autres
Publié: (2024)
par: Copik, Marcin, et autres
Publié: (2024)
Reproducible Cross-border High Performance Computing for Scientific Portals
par: Abarenkov, Kessy, et autres
Publié: (2022)
par: Abarenkov, Kessy, et autres
Publié: (2022)
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
par: Chazapis, Antony, et autres
Publié: (2024)
par: Chazapis, Antony, et autres
Publié: (2024)
Clock Distribution with Gradient TRIX
par: Lenzen, Christoph, et autres
Publié: (2023)
par: Lenzen, Christoph, et autres
Publié: (2023)
Blockchain in a box: A portable blockchain network implementation on Raspberry Pi's
par: Piškorec, Matija, et autres
Publié: (2024)
par: Piškorec, Matija, et autres
Publié: (2024)
High Performance Unstructured SpMM Computation Using Tensor Cores
par: Okanovic, Patrik, et autres
Publié: (2024)
par: Okanovic, Patrik, et autres
Publié: (2024)
Documents similaires
-
Nonlinear spectral clustering with C++ GraphBLAS
par: Pasadakis, Dimosthenis, et autres
Publié: (2026) -
Hypersparse Traffic Matrices from Suricata Network Flows using GraphBLAS
par: Houle, Michael, et autres
Publié: (2024) -
Julia GraphBLAS with Nonblocking Execution
par: Costanza, Pascal, et autres
Publié: (2025) -
Faster Distributed Inference-Only Recommender Systems via Bounded Lag Synchronous Collectives
par: Dichev, Kiril, et autres
Publié: (2025) -
HiCR, an Abstract Model for Distributed Heterogeneous Programming
par: Martin, Sergio Miguel, et autres
Publié: (2025)