Algebraic Temporal Blocking for Sparse Iterative Solvers on Multi-Core CPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Alappat, Christie, Thies, Jonas, Hager, Georg, Fehske, Holger, Wellein, Gerhard |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
by: Lacey, Dane C., et al.
Published: (2024)
by: Lacey, Dane C., et al.
Published: (2024)
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
by: Laukemann, Jan, et al.
Published: (2024)
by: Laukemann, Jan, et al.
Published: (2024)
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
by: Laukemann, Jan, et al.
Published: (2023)
by: Laukemann, Jan, et al.
Published: (2023)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
by: Afzal, Ayesha, et al.
Published: (2024)
by: Afzal, Ayesha, et al.
Published: (2024)
Parallel Sparse and Data-Sparse Factorization-based Linear Solvers
by: Li, Xiaoye Sherry, et al.
Published: (2026)
by: Li, Xiaoye Sherry, et al.
Published: (2026)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
by: Afzal, Ayesha, et al.
Published: (2025)
by: Afzal, Ayesha, et al.
Published: (2025)
Two-Stage Block Orthogonalization to Improve Performance of $s$-step GMRES
by: Yamazaki, Ichitaro, et al.
Published: (2024)
by: Yamazaki, Ichitaro, et al.
Published: (2024)
Analytical Performance Estimation during Code Generation on Modern GPUs
by: Ernst, Dominik, et al.
Published: (2022)
by: Ernst, Dominik, et al.
Published: (2022)
Random-sketching Techniques to Enhance the Numerical Stability of Block Orthogonalization Algorithms for s-step GMRES
by: Yamazaki, Ichitaro, et al.
Published: (2025)
by: Yamazaki, Ichitaro, et al.
Published: (2025)
Towards a GPU-Parallelization of the neXtSIM-DG Dynamical Core
by: Jendersie, Robert, et al.
Published: (2024)
by: Jendersie, Robert, et al.
Published: (2024)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
by: Afzal, Ayesha, et al.
Published: (2026)
by: Afzal, Ayesha, et al.
Published: (2026)
On the Challenges of Energy-Efficiency Analysis in HPC Systems: Evaluating Synthetic Benchmarks and Gromacs
by: Machado, Rafael Ravedutti Lucio, et al.
Published: (2025)
by: Machado, Rafael Ravedutti Lucio, et al.
Published: (2025)
DiffPhD: A Unified Differentiable Solver for Projective Heterogeneous Materials in Elastodynamics with Contact-Rich GPU-Acceleration
by: Lai, Shih-Yu, et al.
Published: (2026)
by: Lai, Shih-Yu, et al.
Published: (2026)
A Distributed Block Chebyshev-Davidson Algorithm for Parallel Spectral Clustering
by: Pang, Qiyuan, et al.
Published: (2022)
by: Pang, Qiyuan, et al.
Published: (2022)
A simple GPU implementation of spectral-element methods for solving 3D Poisson type equations on rectangular domains and its applications
by: Liu, Xinyu, et al.
Published: (2023)
by: Liu, Xinyu, et al.
Published: (2023)
Error Analysis of Matrix Multiplication Emulation Using Ozaki-II Scheme
by: Uchino, Yuki, et al.
Published: (2026)
by: Uchino, Yuki, et al.
Published: (2026)
Parallel simulation and adaptive mesh refinement for 3D elastostatic contact mechanics problems between deformable bodies
by: Epalle, Alexandre, et al.
Published: (2025)
by: Epalle, Alexandre, et al.
Published: (2025)
Teaching An Old Dog New Tricks: Porting Legacy Code to Heterogeneous Compute Architectures With Automated Code Translation
by: Nytko, Nicolas, et al.
Published: (2025)
by: Nytko, Nicolas, et al.
Published: (2025)
SUNDIALS Time Integrators for Exascale Applications with Many Independent ODE Systems
by: Balos, Cody J., et al.
Published: (2024)
by: Balos, Cody J., et al.
Published: (2024)
Neural Acceleration of Incomplete Cholesky Preconditioners
by: Booth, Joshua Dennis, et al.
Published: (2024)
by: Booth, Joshua Dennis, et al.
Published: (2024)
On some orthogonalization schemes in Tensor Train format
by: Coulaud, Olivier, et al.
Published: (2022)
by: Coulaud, Olivier, et al.
Published: (2022)
Asymptotic Analysis of a Leader Election Algorithm
by: Lavault, Christian, et al.
Published: (2006)
by: Lavault, Christian, et al.
Published: (2006)
A Parallel in Time Algorithm Based on ParaExp for Optimal Control Problems
by: Kwok, Felix, et al.
Published: (2024)
by: Kwok, Felix, et al.
Published: (2024)
Cucheb: A GPU implementation of the filtered Lanczos procedure
by: Aurentz, Jared L., et al.
Published: (2024)
by: Aurentz, Jared L., et al.
Published: (2024)
RAPTOR: Practical Numerical Profiling of Scientific Applications
by: Hoerold, Faveo, et al.
Published: (2025)
by: Hoerold, Faveo, et al.
Published: (2025)
GPU Accelerated Implicit Kinetic Meshfree Method based on Modified LU-SGS
by: Verma, Mayuri, et al.
Published: (2024)
by: Verma, Mayuri, et al.
Published: (2024)
Residual-Weighted Randomized Jacobi: Sharpened Bounds via Residual Concentration and Asynchronous Extension
by: Coleman, Evan
Published: (2026)
by: Coleman, Evan
Published: (2026)
CG-Kit: Code Generation Toolkit for Performant and Maintainable Variants of Source Code Applied to Flash-X Hydrodynamics Simulations
by: Rudi, Johann, et al.
Published: (2024)
by: Rudi, Johann, et al.
Published: (2024)
Modifying the Asynchronous Jacobi Method for Data Corruption Resilience
by: Vogl, Christopher J., et al.
Published: (2022)
by: Vogl, Christopher J., et al.
Published: (2022)
Fully-Automated Code Generation for Efficient Computation of Sparse Matrix Permanents on GPUs
by: Elbek, Deniz, et al.
Published: (2025)
by: Elbek, Deniz, et al.
Published: (2025)
Efficient and scalable atmospheric dynamics simulations using non-conforming meshes
by: Orlando, Giuseppe, et al.
Published: (2024)
by: Orlando, Giuseppe, et al.
Published: (2024)
Real-time Bayesian inference at extreme scale: A digital twin for tsunami early warning applied to the Cascadia subduction zone
by: Henneking, Stefan, et al.
Published: (2025)
by: Henneking, Stefan, et al.
Published: (2025)
Improving the scalability of a high-order atmospheric dynamics solver based on the deal.II library
by: Orlando, Giuseppe, et al.
Published: (2025)
by: Orlando, Giuseppe, et al.
Published: (2025)
A High Performance GPU CountSketch Implementation and Its Application to Multisketching and Least Squares Problems
by: Higgins, Andrew J., et al.
Published: (2025)
by: Higgins, Andrew J., et al.
Published: (2025)
Randomized algorithms for distributed computation of principal component analysis and singular value decomposition
by: Li, Huamin, et al.
Published: (2016)
by: Li, Huamin, et al.
Published: (2016)
GPU-Parallelizable Randomized Sketch-and-Precondition for Linear Regression using Sparse Sign Sketches
by: Chen, Tyler, et al.
Published: (2025)
by: Chen, Tyler, et al.
Published: (2025)
A Domain Decomposition-based Solver for Acoustic Wave propagation in Two-Dimensional Random Media
by: Vasudevan, Sudhi Sharma Padillath
Published: (2025)
by: Vasudevan, Sudhi Sharma Padillath
Published: (2025)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
by: Wolfson-Pou, Jordi, et al.
Published: (2025)
by: Wolfson-Pou, Jordi, et al.
Published: (2025)
SUperman: Efficient Permanent Computation on GPUs
by: Elbek, Deniz, et al.
Published: (2025)
by: Elbek, Deniz, et al.
Published: (2025)
TTrace: Lightweight Error Checking and Diagnosis for Distributed Training
by: Jiang, Haitian, et al.
Published: (2025)
by: Jiang, Haitian, et al.
Published: (2025)
Similar Items
-
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
by: Lacey, Dane C., et al.
Published: (2024) -
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
by: Laukemann, Jan, et al.
Published: (2024) -
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
by: Laukemann, Jan, et al.
Published: (2023) -
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
by: Afzal, Ayesha, et al.
Published: (2024) -
Parallel Sparse and Data-Sparse Factorization-based Linear Solvers
by: Li, Xiaoye Sherry, et al.
Published: (2026)