Mixed-Precision Performance Portability of FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices
Fuente:
arXiv
Saved in:
| Main Authors: | Venkat, Sreeram, Swirydowicz, Kasia, Wolfe, Noah, Ghattas, Omar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
High-performance matrix-free unfitted finite element operator evaluation
by: Bergbauer, Maximilian, et al.
Published: (2024)
by: Bergbauer, Maximilian, et al.
Published: (2024)
Matrix-Free Evaluation of High-Order Shifted Boundary Finite Element Operators
by: Wichrowski, Michał
Published: (2025)
by: Wichrowski, Michał
Published: (2025)
Scalable Multilevel Monte Carlo Methods Exploiting Parallel Redistribution on Coarse Levels
by: Fairbanks, Hillary R., et al.
Published: (2024)
by: Fairbanks, Hillary R., et al.
Published: (2024)
Any nonincreasing convergence curves are simultaneously possible for GMRES and weighted GMRES, as well as for left and right preconditioned GMRES
by: Matalon, Pierre, et al.
Published: (2025)
by: Matalon, Pierre, et al.
Published: (2025)
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
by: Kriemann, Ronald
Published: (2024)
by: Kriemann, Ronald
Published: (2024)
Stochastic trace estimation for parameter-dependent matrices applied to spectral density approximation
by: Matti, Fabio, et al.
Published: (2025)
by: Matti, Fabio, et al.
Published: (2025)
Mixed Precision Orthogonalization-Free Projection Methods for Eigenvalue and Singular Value Problems
by: Xu, Tianshi, et al.
Published: (2025)
by: Xu, Tianshi, et al.
Published: (2025)
Parallel-in-time Multilevel Krylov Methods: A Prototype
by: Erlangga, Yogi A.
Published: (2023)
by: Erlangga, Yogi A.
Published: (2023)
Robust Blockwise Random Pivoting: Fast and Accurate Adaptive Interpolative Decomposition
by: Dong, Yijun, et al.
Published: (2023)
by: Dong, Yijun, et al.
Published: (2023)
A resource-efficient model for deep kernel learning
by: D'Amore, Luisa
Published: (2024)
by: D'Amore, Luisa
Published: (2024)
ML-Based Optimum Number of CUDA Streams for the GPU Implementation of the Tridiagonal Partition Method
by: Veneva, Milena, et al.
Published: (2025)
by: Veneva, Milena, et al.
Published: (2025)
ML-Based Optimum Sub-system Size Heuristic for the GPU Implementation of the Tridiagonal Partition Method
by: Veneva, Milena
Published: (2025)
by: Veneva, Milena
Published: (2025)
Iterative Methods in GPU-Resident Linear Solvers for Nonlinear Constrained Optimization
by: Świrydowicz, Kasia, et al.
Published: (2024)
by: Świrydowicz, Kasia, et al.
Published: (2024)
Subspace-constrained randomized coordinate descent for linear systems with good low-rank matrix approximations
by: Lok, Jackie, et al.
Published: (2025)
by: Lok, Jackie, et al.
Published: (2025)
Space-time parallel scaling of Parareal with a physics-informed Fourier Neural Operator coarse propagator applied to the Black-Scholes equation
by: Ibrahim, Abdul Qadir, et al.
Published: (2024)
by: Ibrahim, Abdul Qadir, et al.
Published: (2024)
Low-Memory Numerical Certification
by: Breiding, Paul, et al.
Published: (2026)
by: Breiding, Paul, et al.
Published: (2026)
Communication efficient application of sequences of planar rotations to a matrix
by: Steel, Thijs, et al.
Published: (2024)
by: Steel, Thijs, et al.
Published: (2024)
A multigrid method for CutFEM and its implementation on GPU
by: Cui, Cu, et al.
Published: (2025)
by: Cui, Cu, et al.
Published: (2025)
M2L Translation Operators for Kernel Independent Fast Multipole Methods on Modern Architectures
by: Kailasa, Srinath, et al.
Published: (2024)
by: Kailasa, Srinath, et al.
Published: (2024)
Communication-reduced Conjugate Gradient Variants for GPU-accelerated Clusters
by: Bernaschi, Massimo, et al.
Published: (2025)
by: Bernaschi, Massimo, et al.
Published: (2025)
Memory- and compute-optimized geometric multigrid GMGPolar for curvilinear coordinate representations -- Applications to fusion plasma
by: Litz, Julian, et al.
Published: (2025)
by: Litz, Julian, et al.
Published: (2025)
Distributed Parallel Structure-Aware Presolving for Arrowhead Linear Programs
by: Kempke, Nils-Christian, et al.
Published: (2026)
by: Kempke, Nils-Christian, et al.
Published: (2026)
Space-Time Block Preconditioning for Incompressible Resistive Magnetohydrodynamics
by: Danieli, Federico, et al.
Published: (2023)
by: Danieli, Federico, et al.
Published: (2023)
Bare-Metal Tensor Virtualization: Overcoming the Memory Wall in Edge-AI Inference on ARM64
by: Kilictas, Bugra, et al.
Published: (2026)
by: Kilictas, Bugra, et al.
Published: (2026)
Massively Parallel Reductions in Multivariate Polynomial Systems: Bridging the Symbolic Preprocessing Gap on GPGPU Architectures
by: Gokavarapu, Chandrasekhar
Published: (2026)
by: Gokavarapu, Chandrasekhar
Published: (2026)
Hybrid hierarchical matrices with adaptive mixed precision storage
by: Khan, Ritesh, et al.
Published: (2026)
by: Khan, Ritesh, et al.
Published: (2026)
Mixed precision thin SVD algorithms based on the Gram matrix
by: Carson, Erin, et al.
Published: (2026)
by: Carson, Erin, et al.
Published: (2026)
MIRGE: An Array-Based Computational Framework for Scientific Computing
by: Diener, Matthias, et al.
Published: (2025)
by: Diener, Matthias, et al.
Published: (2025)
An $O(\log N)$ Monte Carlo method for periodic Coulomb systems
by: Gao, Xuanzhao, et al.
Published: (2026)
by: Gao, Xuanzhao, et al.
Published: (2026)
Space-time parallel iterative solvers for the integration of parabolic problems
by: Arrarás, Andrés, et al.
Published: (2025)
by: Arrarás, Andrés, et al.
Published: (2025)
A Parareal Algorithm with Low-Rank Coarse Solvers
by: Gander, Martin J., et al.
Published: (2025)
by: Gander, Martin J., et al.
Published: (2025)
Parallelization Strategies for the Randomized Kaczmarz Algorithm on Large-Scale Dense Systems
by: Ferreira, Inês, et al.
Published: (2024)
by: Ferreira, Inês, et al.
Published: (2024)
Surrogate-based Autotuning for Randomized Sketching Algorithms in Regression Problems
by: Cho, Younghyun, et al.
Published: (2023)
by: Cho, Younghyun, et al.
Published: (2023)
A Virtual Processor brings back the Free Lunch
by: Kutschbach, Haymo
Published: (2026)
by: Kutschbach, Haymo
Published: (2026)
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
Hybrid metapopulation agent-based epidemiological models for efficient insight on the individual scale: a contribution to green computing
by: Bicker, Julia, et al.
Published: (2024)
by: Bicker, Julia, et al.
Published: (2024)
Parallel Gauss-Jordan Elimination and System Reduction for Efficient Circuit Simulation
by: Noveski, Filip, et al.
Published: (2026)
by: Noveski, Filip, et al.
Published: (2026)
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
RandNet-Parareal: a time-parallel PDE solver using Random Neural Networks
by: Gattiglio, Guglielmo, et al.
Published: (2024)
by: Gattiglio, Guglielmo, et al.
Published: (2024)
Fast GPU Linear Algebra via Compile Time Expression Fusion
by: Curtin, Ryan R., et al.
Published: (2026)
by: Curtin, Ryan R., et al.
Published: (2026)
Similar Items
-
High-performance matrix-free unfitted finite element operator evaluation
by: Bergbauer, Maximilian, et al.
Published: (2024) -
Matrix-Free Evaluation of High-Order Shifted Boundary Finite Element Operators
by: Wichrowski, Michał
Published: (2025) -
Scalable Multilevel Monte Carlo Methods Exploiting Parallel Redistribution on Coarse Levels
by: Fairbanks, Hillary R., et al.
Published: (2024) -
Any nonincreasing convergence curves are simultaneously possible for GMRES and weighted GMRES, as well as for left and right preconditioned GMRES
by: Matalon, Pierre, et al.
Published: (2025) -
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
by: Kriemann, Ronald
Published: (2024)