Towards a Higher Roofline for Matrix-Vector Multiplication in Matrix-Free HOSFEM
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Zijian, Sun, Qiao, Zhang, Tiangong, Li, Huiyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Test for FLOPs as a Discriminant for Linear Algebra Algorithms
by: Sankaran, Aravind, et al.
Published: (2022)
by: Sankaran, Aravind, et al.
Published: (2022)
RAO-SS: A Prototype of Run-time Auto-tuning Facility for Sparse Direct Solvers
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
Faster Base64 Encoding and Decoding Using AVX2 Instructions
by: Muła, Wojciech, et al.
Published: (2017)
by: Muła, Wojciech, et al.
Published: (2017)
Code Generation for Near-Roofline Finite Element Actions on GPUs from Symbolic Variational Forms
by: Kulkarni, Kaushik, et al.
Published: (2025)
by: Kulkarni, Kaushik, et al.
Published: (2025)
rcpptimer: Rcpp Tic-Toc Timer with OpenMP Support
by: Berrisch, Jonathan
Published: (2025)
by: Berrisch, Jonathan
Published: (2025)
Matrix-Free Evaluation Strategies for Continuous and Discontinuous Galerkin Discretizations on Unstructured Tetrahedral Grids
by: Still, Dominik, et al.
Published: (2025)
by: Still, Dominik, et al.
Published: (2025)
Building an Accelerated OpenFOAM Proof-of-Concept Application using Modern C++
by: Malenza, Giulio, et al.
Published: (2025)
by: Malenza, Giulio, et al.
Published: (2025)
DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
by: Mukunoki, Daichi
Published: (2025)
by: Mukunoki, Daichi
Published: (2025)
Annotation-guided AoS-to-SoA conversions and GPU offloading with data views in C++
by: Radtke, Pawel K., et al.
Published: (2025)
by: Radtke, Pawel K., et al.
Published: (2025)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
by: Li, Junjie
Published: (2024)
by: Li, Junjie
Published: (2024)
Toward Capturing Genetic Epistasis From Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression
by: Ltaief, Hatem, et al.
Published: (2024)
by: Ltaief, Hatem, et al.
Published: (2024)
Giga-scale Kernel Matrix Vector Multiplication on GPU
by: Hu, Robert, et al.
Published: (2022)
by: Hu, Robert, et al.
Published: (2022)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
by: Tuteja, Keshvi, et al.
Published: (2025)
by: Tuteja, Keshvi, et al.
Published: (2025)
On the energy efficiency of sparse matrix computations on multi-GPU clusters
by: Bernaschi, Massimo, et al.
Published: (2025)
by: Bernaschi, Massimo, et al.
Published: (2025)
NApy: Efficient Statistics in Python for Large-Scale Heterogeneous Data with Enhanced Support for Missing Data
by: Woller, Fabian, et al.
Published: (2025)
by: Woller, Fabian, et al.
Published: (2025)
Performance measurements of modern Fortran MPI applications with Score-P
by: Corbin, Gregor
Published: (2025)
by: Corbin, Gregor
Published: (2025)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
Automated MPI-X code generation for scalable finite-difference solvers
by: Bisbas, George, et al.
Published: (2023)
by: Bisbas, George, et al.
Published: (2023)
A Communication Avoiding and Reducing Algorithm for Symmetric Eigenproblem for Very Small Matrices
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
by: Panova, Elena, et al.
Published: (2022)
by: Panova, Elena, et al.
Published: (2022)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
by: Tu, Jiqun, et al.
Published: (2026)
by: Tu, Jiqun, et al.
Published: (2026)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
Verification Challenges in Sparse Matrix Vector Multiplication in High Performance Computing: Part I
by: Zhang, Junchao
Published: (2025)
by: Zhang, Junchao
Published: (2025)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
by: Asudeh, Omid, et al.
Published: (2025)
by: Asudeh, Omid, et al.
Published: (2025)
Improving the Graph Challenge Reference Implementation
by: Voloshchuk, Inna, et al.
Published: (2026)
by: Voloshchuk, Inna, et al.
Published: (2026)
Matrix-Free Finite Volume Kernels on a Dataflow Architecture
by: Sai, Ryuichi, et al.
Published: (2024)
by: Sai, Ryuichi, et al.
Published: (2024)
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
by: Zhu, Honglin, et al.
Published: (2026)
by: Zhu, Honglin, et al.
Published: (2026)
A Performance Portable Matrix Free Dense MTTKRP in GenTen
by: Kosmacher, Gabriel, et al.
Published: (2025)
by: Kosmacher, Gabriel, et al.
Published: (2025)
GoldbachGPU: An Open Source GPU-Accelerated Framework for Verification of Goldbach's Conjecture
by: Llorente-Saguer, Isaac
Published: (2026)
by: Llorente-Saguer, Isaac
Published: (2026)
Hierarchical Recursive Precision for Accelerating Symmetric Linear Solves on MXUs
by: Carrica, Vicki, et al.
Published: (2026)
by: Carrica, Vicki, et al.
Published: (2026)
Anonymized Network Sensing Graph Challenge
by: Jananthan, Hayden, et al.
Published: (2024)
by: Jananthan, Hayden, et al.
Published: (2024)
Deriving Algorithms for Triangular Tridiagonalization a Skew-Symmetric Matrix
by: van de Geijn, Robert, et al.
Published: (2023)
by: van de Geijn, Robert, et al.
Published: (2023)
Towards Assessing Spread in Sets of Software Architecture Designs
by: Cortellessa, Vittorio, et al.
Published: (2024)
by: Cortellessa, Vittorio, et al.
Published: (2024)
Shifting the Sweet Spot: High-Performance Matrix-Free Method for High-Order Elasticity
by: Chang, Dali, et al.
Published: (2026)
by: Chang, Dali, et al.
Published: (2026)
TypedMatrices.jl: An Extensible and Type-Based Matrix Collection for Julia
by: Zhang, Anzhi, et al.
Published: (2025)
by: Zhang, Anzhi, et al.
Published: (2025)
High Performance Matrix Multiplication
by: Davis, Ethan
Published: (2025)
by: Davis, Ethan
Published: (2025)
Acceleration of Tensor-Product Operations with Tensor Cores
by: Cui, Cu
Published: (2024)
by: Cui, Cu
Published: (2024)
LLM-Vectorizer: LLM-based Verified Loop Vectorizer
by: Taneja, Jubi, et al.
Published: (2024)
by: Taneja, Jubi, et al.
Published: (2024)
Performance Analysis of Effective Symbolic Methods for Solving Band Matrix SLAEs
by: Veneva, Milena, et al.
Published: (2019)
by: Veneva, Milena, et al.
Published: (2019)
Similar Items
-
A Test for FLOPs as a Discriminant for Linear Algebra Algorithms
by: Sankaran, Aravind, et al.
Published: (2022) -
RAO-SS: A Prototype of Run-time Auto-tuning Facility for Sparse Direct Solvers
by: Katagiri, Takahiro, et al.
Published: (2024) -
Faster Base64 Encoding and Decoding Using AVX2 Instructions
by: Muła, Wojciech, et al.
Published: (2017) -
Code Generation for Near-Roofline Finite Element Actions on GPUs from Symbolic Variational Forms
by: Kulkarni, Kaushik, et al.
Published: (2025) -
rcpptimer: Rcpp Tic-Toc Timer with OpenMP Support
by: Berrisch, Jonathan
Published: (2025)