Verification Challenges in Sparse Matrix Vector Multiplication in High Performance Computing: Part I
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Zhang, Junchao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
von: Li, Yifan, et al.
Veröffentlicht: (2026)
von: Li, Yifan, et al.
Veröffentlicht: (2026)
Reusable Formal Verification of DAG-based Consensus Protocols
von: Bertrand, Nathalie, et al.
Veröffentlicht: (2024)
von: Bertrand, Nathalie, et al.
Veröffentlicht: (2024)
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
von: Zhu, Honglin, et al.
Veröffentlicht: (2026)
von: Zhu, Honglin, et al.
Veröffentlicht: (2026)
Modelling the Raft Distributed Consensus Protocol in mCRL2
von: Bora, Parth, et al.
Veröffentlicht: (2024)
von: Bora, Parth, et al.
Veröffentlicht: (2024)
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
von: Ringoot, Evelyne, et al.
Veröffentlicht: (2025)
von: Ringoot, Evelyne, et al.
Veröffentlicht: (2025)
High-Performance Star-M SVD for Big Data Compression
von: Hussain, Md Taufique, et al.
Veröffentlicht: (2026)
von: Hussain, Md Taufique, et al.
Veröffentlicht: (2026)
Complexity of Verification and Synthesis of Threshold Automata
von: Balasubramanian, A. R., et al.
Veröffentlicht: (2020)
von: Balasubramanian, A. R., et al.
Veröffentlicht: (2020)
Distributed-memory Algorithms for Sparse Matrix Permutation, Extraction, and Assignment
von: Hassani, Elaheh, et al.
Veröffentlicht: (2025)
von: Hassani, Elaheh, et al.
Veröffentlicht: (2025)
Interactive Safety Verification of Distributed Protocols by Inductive Proof Decomposition
von: Schultz, William, et al.
Veröffentlicht: (2024)
von: Schultz, William, et al.
Veröffentlicht: (2024)
On the Challenges of Energy-Efficiency Analysis in HPC Systems: Evaluating Synthetic Benchmarks and Gromacs
von: Machado, Rafael Ravedutti Lucio, et al.
Veröffentlicht: (2025)
von: Machado, Rafael Ravedutti Lucio, et al.
Veröffentlicht: (2025)
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
von: Carrica, Vicki, et al.
Veröffentlicht: (2025)
von: Carrica, Vicki, et al.
Veröffentlicht: (2025)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
Continuous reasoning for adaptive container image distribution in the cloud-edge continuum
von: Azzolini, Damiano, et al.
Veröffentlicht: (2024)
von: Azzolini, Damiano, et al.
Veröffentlicht: (2024)
Parallel Sparse and Data-Sparse Factorization-based Linear Solvers
von: Li, Xiaoye Sherry, et al.
Veröffentlicht: (2026)
von: Li, Xiaoye Sherry, et al.
Veröffentlicht: (2026)
Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
von: Wang, Hansheng, et al.
Veröffentlicht: (2025)
von: Wang, Hansheng, et al.
Veröffentlicht: (2025)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
von: Tuteja, Keshvi, et al.
Veröffentlicht: (2025)
von: Tuteja, Keshvi, et al.
Veröffentlicht: (2025)
Performance measurements of modern Fortran MPI applications with Score-P
von: Corbin, Gregor
Veröffentlicht: (2025)
von: Corbin, Gregor
Veröffentlicht: (2025)
The Ubiquitous Sparse Matrix-Matrix Products
von: Buluç, Aydın
Veröffentlicht: (2025)
von: Buluç, Aydın
Veröffentlicht: (2025)
Enabling mixed-precision in spectral element codes
von: Chen, Yanxiang, et al.
Veröffentlicht: (2025)
von: Chen, Yanxiang, et al.
Veröffentlicht: (2025)
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
von: Ringoot, Evelyne, et al.
Veröffentlicht: (2025)
von: Ringoot, Evelyne, et al.
Veröffentlicht: (2025)
Integrating Odeint Time Stepping into OpenFPM for Distributed and GPU Accelerated Numerical Solvers
von: Singh, Abhinav, et al.
Veröffentlicht: (2023)
von: Singh, Abhinav, et al.
Veröffentlicht: (2023)
A shared compilation stack for distributed-memory parallelism in stencil DSLs
von: Bisbas, George, et al.
Veröffentlicht: (2024)
von: Bisbas, George, et al.
Veröffentlicht: (2024)
Robustness and Accuracy in Pipelined Bi-Conjugate Gradient Stabilized Method: A Comparative Study
von: Havdiak, Mykhailo, et al.
Veröffentlicht: (2024)
von: Havdiak, Mykhailo, et al.
Veröffentlicht: (2024)
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
von: Bellavita, Julian, et al.
Veröffentlicht: (2026)
von: Bellavita, Julian, et al.
Veröffentlicht: (2026)
Efficient N-to-M Checkpointing Algorithm for Finite Element Simulations
von: Ham, David A., et al.
Veröffentlicht: (2024)
von: Ham, David A., et al.
Veröffentlicht: (2024)
A new open source framework for multiscale modeling of fibrous materials on heterogeneous supercomputers
von: Merson, Jacob, et al.
Veröffentlicht: (2023)
von: Merson, Jacob, et al.
Veröffentlicht: (2023)
Enabling MPI communication within Numba/LLVM JIT-compiled Python code using numba-mpi v1.0
von: Derlatka, Kacper, et al.
Veröffentlicht: (2024)
von: Derlatka, Kacper, et al.
Veröffentlicht: (2024)
SYCL compute kernels for ExaHyPE
von: Loi, Chung Ming, et al.
Veröffentlicht: (2023)
von: Loi, Chung Ming, et al.
Veröffentlicht: (2023)
Verification of Population Protocols with Unordered Data
von: van Bergerem, Steffen, et al.
Veröffentlicht: (2024)
von: van Bergerem, Steffen, et al.
Veröffentlicht: (2024)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
GPU Implementations for Midsize Integer Addition and Multiplication
von: Oancea, Cosmin E., et al.
Veröffentlicht: (2024)
von: Oancea, Cosmin E., et al.
Veröffentlicht: (2024)
AMReX: Block-Structured Adaptive Mesh Refinement for Multiphysics Applications
von: Zhang, Weiqun, et al.
Veröffentlicht: (2020)
von: Zhang, Weiqun, et al.
Veröffentlicht: (2020)
Towards a Formal Verification of Secure Vehicle Software Updates
von: Hagen, Martin Slind, et al.
Veröffentlicht: (2025)
von: Hagen, Martin Slind, et al.
Veröffentlicht: (2025)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
von: Li, Junjie
Veröffentlicht: (2024)
von: Li, Junjie
Veröffentlicht: (2024)
Enabling mixed-precision with the help of tools: A Nekbone case study
von: Chen, Yanxiang, et al.
Veröffentlicht: (2024)
von: Chen, Yanxiang, et al.
Veröffentlicht: (2024)
Proceedings 18th Interaction and Concurrency Experience
von: Aubert, Clément, et al.
Veröffentlicht: (2025)
von: Aubert, Clément, et al.
Veröffentlicht: (2025)
Application Placement with Constraint Relaxation
von: Azzolini, Damiano, et al.
Veröffentlicht: (2025)
von: Azzolini, Damiano, et al.
Veröffentlicht: (2025)
A categorical and logical framework for iterated protocols
von: Goubault, Eric, et al.
Veröffentlicht: (2025)
von: Goubault, Eric, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
von: Li, Yifan, et al.
Veröffentlicht: (2026) -
Reusable Formal Verification of DAG-based Consensus Protocols
von: Bertrand, Nathalie, et al.
Veröffentlicht: (2024) -
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
von: Zhu, Honglin, et al.
Veröffentlicht: (2026) -
Modelling the Raft Distributed Consensus Protocol in mCRL2
von: Bora, Parth, et al.
Veröffentlicht: (2024) -
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)