Fast GPU Linear Algebra via Compile Time Expression Fusion
Fuente:
arXiv
Saved in:
| Main Authors: | Curtin, Ryan R., Edel, Marcus, Sanderson, Conrad |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Armadillo: An Efficient Framework for Numerical Linear Algebra
by: Sanderson, Conrad, et al.
Published: (2025)
by: Sanderson, Conrad, et al.
Published: (2025)
Bandicoot: A Templated C++ Library for GPU Linear Algebra
by: Curtin, Ryan R., et al.
Published: (2025)
by: Curtin, Ryan R., et al.
Published: (2025)
Mixed Precision FGMRES-Based Iterative Refinement for Weighted Least Squares
by: Carson, Erin, et al.
Published: (2024)
by: Carson, Erin, et al.
Published: (2024)
Fast Evaluation of Truncated Neumann Series by Low-Product Radix Kernels
by: Sao, Piyush
Published: (2026)
by: Sao, Piyush
Published: (2026)
Inexact Gauss Seidel and Coarse Solvers for AMG and s-step CG
by: Thomas, Stephen, et al.
Published: (2025)
by: Thomas, Stephen, et al.
Published: (2025)
Computing the spectrum and pseudospectrum of infinite-volume operators from local patches
by: Hege, Paul, et al.
Published: (2024)
by: Hege, Paul, et al.
Published: (2024)
Entrywise Approximation for Matrix Inversion and Linear Systems
by: Ghadiri, Mehrdad, et al.
Published: (2025)
by: Ghadiri, Mehrdad, et al.
Published: (2025)
Data-Adaptive Low-Rank Sparse Subspace Clustering
by: Kopriva, Ivica
Published: (2025)
by: Kopriva, Ivica
Published: (2025)
Scalable s-step Preconditioned Conjugate Gradient with Chebyshev Basis and Gauss-Seidel Gram Solve
by: D'Ambra, Pasqua, et al.
Published: (2026)
by: D'Ambra, Pasqua, et al.
Published: (2026)
Fast Gaussian process inference by exact Matérn kernel decomposition
by: Langrené, Nicolas, et al.
Published: (2025)
by: Langrené, Nicolas, et al.
Published: (2025)
What is a POLYNOMIAL-TIME Computable L2-Function?
by: Bacho, Aras, et al.
Published: (2026)
by: Bacho, Aras, et al.
Published: (2026)
Symbolic Algorithm for Solving SLAEs with Multi-Diagonal Coefficient Matrices
by: Veneva, Milena
Published: (2024)
by: Veneva, Milena
Published: (2024)
Multiprecision computations with Schwarz methods
by: Outrata, Michal, et al.
Published: (2025)
by: Outrata, Michal, et al.
Published: (2025)
Reorthogonalized Pythagorean variants of block classical Gram-Schmidt
by: Carson, Erin, et al.
Published: (2024)
by: Carson, Erin, et al.
Published: (2024)
On the loss of orthogonality in low-synchronization variants of reorthogonalized block classical Gram-Schmidt
by: Carson, Erin, et al.
Published: (2024)
by: Carson, Erin, et al.
Published: (2024)
Recursive vectorized computation of the vector $p$-norm
by: Novaković, Vedran
Published: (2025)
by: Novaković, Vedran
Published: (2025)
Inexact and primal multilevel FETI-DP methods: a multilevel extension and interplay with BDDC
by: Sousedík, Bedřich
Published: (2022)
by: Sousedík, Bedřich
Published: (2022)
Space-Time Trade-off in Integer Linear Scaling Rounded to the Nearest Integer through Multiplicative and Additive Decomposition
by: Kim, Kyeong Soo
Published: (2026)
by: Kim, Kyeong Soo
Published: (2026)
Trilinos: Enabling Scientific Computing Across Diverse Hardware Architectures at Scale
by: Mayr, Matthias, et al.
Published: (2025)
by: Mayr, Matthias, et al.
Published: (2025)
Parallelization and scalability analysis of inverse factorization using the Chunks and Tasks programming model
by: Artemov, Anton G., et al.
Published: (2019)
by: Artemov, Anton G., et al.
Published: (2019)
Direct Low-Dose CT Image Reconstruction on GPU using Out-Of-Core: Precision and Quality Study
by: Chillarón, M., et al.
Published: (2024)
by: Chillarón, M., et al.
Published: (2024)
Anderson acceleration with approximate calculations: applications to scientific computing
by: Pasini, Massimiliano Lupo, et al.
Published: (2022)
by: Pasini, Massimiliano Lupo, et al.
Published: (2022)
Gradient Coding with Iterative Block Leverage Score Sampling
by: Charalambides, Neophytos, et al.
Published: (2023)
by: Charalambides, Neophytos, et al.
Published: (2023)
Analysis of Floating-Point Matrix Multiplication Computed via Integer Arithmetic
by: Abdelfattah, Ahmad, et al.
Published: (2025)
by: Abdelfattah, Ahmad, et al.
Published: (2025)
Mixed precision thin SVD algorithms based on the Gram matrix
by: Carson, Erin, et al.
Published: (2026)
by: Carson, Erin, et al.
Published: (2026)
Subspace gradient descent method for linear tensor equations
by: Iannacito, Martina, et al.
Published: (2026)
by: Iannacito, Martina, et al.
Published: (2026)
Scalable Preconditioners for the Pseudo-4D DFN Lithium-ion Battery Model
by: Roy, Thomas, et al.
Published: (2026)
by: Roy, Thomas, et al.
Published: (2026)
The complexity of accurate floating point computation
by: Demmel, James
Published: (2003)
by: Demmel, James
Published: (2003)
Semantics, Specification Logic, and Hoare Logic of Exact Real Computation
by: Park, Sewon, et al.
Published: (2016)
by: Park, Sewon, et al.
Published: (2016)
Hardware Trends Impacting Floating-Point Computations In Scientific Applications
by: Dongarra, Jack, et al.
Published: (2024)
by: Dongarra, Jack, et al.
Published: (2024)
Multigrid with Linear Storage Complexity
by: Bauer, Daniel, et al.
Published: (2025)
by: Bauer, Daniel, et al.
Published: (2025)
RANDSMAPs: Random-Feature/multi-Scale Neural Decoders with Mass Preservation
by: Patsatzis, Dimitrios G., et al.
Published: (2026)
by: Patsatzis, Dimitrios G., et al.
Published: (2026)
The ensmallen library for flexible numerical optimization
by: Curtin, Ryan R., et al.
Published: (2021)
by: Curtin, Ryan R., et al.
Published: (2021)
Communication-reduced Conjugate Gradient Variants for GPU-accelerated Clusters
by: Bernaschi, Massimo, et al.
Published: (2025)
by: Bernaschi, Massimo, et al.
Published: (2025)
Arithmetical enhancements of the Kogbetliantz method for the SVD of order two
by: Novaković, Vedran
Published: (2024)
by: Novaković, Vedran
Published: (2024)
Mixed-Precision Performance Portability of FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices
by: Venkat, Sreeram, et al.
Published: (2025)
by: Venkat, Sreeram, et al.
Published: (2025)
A stable one-synchronization variant of reorthogonalized block classical Gram--Schmidt
by: Carson, Erin, et al.
Published: (2024)
by: Carson, Erin, et al.
Published: (2024)
Forward and backward error bounds for a mixed precision preconditioned conjugate gradient algorithm
by: Bake, Thomas, et al.
Published: (2025)
by: Bake, Thomas, et al.
Published: (2025)
Do Inner Greenland's Melt Rate Dynamics Approach Coastal Ones?
by: Heßler, Martin, et al.
Published: (2024)
by: Heßler, Martin, et al.
Published: (2024)
A Parallel-in-Time Combination Method for Parabolic Problems
by: Griebel, Michael, et al.
Published: (2025)
by: Griebel, Michael, et al.
Published: (2025)
Similar Items
-
Armadillo: An Efficient Framework for Numerical Linear Algebra
by: Sanderson, Conrad, et al.
Published: (2025) -
Bandicoot: A Templated C++ Library for GPU Linear Algebra
by: Curtin, Ryan R., et al.
Published: (2025) -
Mixed Precision FGMRES-Based Iterative Refinement for Weighted Least Squares
by: Carson, Erin, et al.
Published: (2024) -
Fast Evaluation of Truncated Neumann Series by Low-Product Radix Kernels
by: Sao, Piyush
Published: (2026) -
Inexact Gauss Seidel and Coarse Solvers for AMG and s-step CG
by: Thomas, Stephen, et al.
Published: (2025)