Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
Fuente:
arXiv
Saved in:
| Main Author: | Kriemann, Ronald |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Matrix-Free Evaluation of High-Order Shifted Boundary Finite Element Operators
by: Wichrowski, Michał
Published: (2025)
by: Wichrowski, Michał
Published: (2025)
Stochastic trace estimation for parameter-dependent matrices applied to spectral density approximation
by: Matti, Fabio, et al.
Published: (2025)
by: Matti, Fabio, et al.
Published: (2025)
Parallelization Strategies for the Randomized Kaczmarz Algorithm on Large-Scale Dense Systems
by: Ferreira, Inês, et al.
Published: (2024)
by: Ferreira, Inês, et al.
Published: (2024)
Massively Parallel Reductions in Multivariate Polynomial Systems: Bridging the Symbolic Preprocessing Gap on GPGPU Architectures
by: Gokavarapu, Chandrasekhar
Published: (2026)
by: Gokavarapu, Chandrasekhar
Published: (2026)
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
Mixed-Precision Performance Portability of FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices
by: Venkat, Sreeram, et al.
Published: (2025)
by: Venkat, Sreeram, et al.
Published: (2025)
High-performance matrix-free unfitted finite element operator evaluation
by: Bergbauer, Maximilian, et al.
Published: (2024)
by: Bergbauer, Maximilian, et al.
Published: (2024)
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
Exploiting nested task-parallelism in the $\mathcal{H}-LU$ factorization
by: Carratalá-Sáez, Rocío, et al.
Published: (2019)
by: Carratalá-Sáez, Rocío, et al.
Published: (2019)
Distributed Hybrid Sketching for $\ell_2$-Embeddings
by: Charalambides, Neophytos, et al.
Published: (2024)
by: Charalambides, Neophytos, et al.
Published: (2024)
CLAIRE: Scalable GPU-Accelerated Algorithms for Diffeomorphic Image Registration in 3D
by: Mang, Andreas
Published: (2024)
by: Mang, Andreas
Published: (2024)
ML-Based Optimum Number of CUDA Streams for the GPU Implementation of the Tridiagonal Partition Method
by: Veneva, Milena, et al.
Published: (2025)
by: Veneva, Milena, et al.
Published: (2025)
ML-Based Optimum Sub-system Size Heuristic for the GPU Implementation of the Tridiagonal Partition Method
by: Veneva, Milena
Published: (2025)
by: Veneva, Milena
Published: (2025)
Hybrid hierarchical matrices with adaptive mixed precision storage
by: Khan, Ritesh, et al.
Published: (2026)
by: Khan, Ritesh, et al.
Published: (2026)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
by: Zhao, Haisha, et al.
Published: (2025)
by: Zhao, Haisha, et al.
Published: (2025)
Scalable Multilevel Monte Carlo Methods Exploiting Parallel Redistribution on Coarse Levels
by: Fairbanks, Hillary R., et al.
Published: (2024)
by: Fairbanks, Hillary R., et al.
Published: (2024)
Any nonincreasing convergence curves are simultaneously possible for GMRES and weighted GMRES, as well as for left and right preconditioned GMRES
by: Matalon, Pierre, et al.
Published: (2025)
by: Matalon, Pierre, et al.
Published: (2025)
Mixed Precision Orthogonalization-Free Projection Methods for Eigenvalue and Singular Value Problems
by: Xu, Tianshi, et al.
Published: (2025)
by: Xu, Tianshi, et al.
Published: (2025)
Algorithms for Parallel Shared-Memory Sparse Matrix-Vector Multiplication on Unstructured Matrices
by: Bergmans, Kobe, et al.
Published: (2025)
by: Bergmans, Kobe, et al.
Published: (2025)
Distributed Tomographic Reconstruction with Quantization
by: Miao, Runxuan, et al.
Published: (2024)
by: Miao, Runxuan, et al.
Published: (2024)
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Accelerating State-Vector Quantum Simulation on Integrated GPUs via Cache Locality Optimization: A Cross-Architecture Evaluation
by: Thomaz, Gabriel Fernandes, et al.
Published: (2026)
by: Thomaz, Gabriel Fernandes, et al.
Published: (2026)
Population Protocols Revisited: Parity and Beyond
by: Gąsieniec, Leszek, et al.
Published: (2025)
by: Gąsieniec, Leszek, et al.
Published: (2025)
Subspace-constrained randomized coordinate descent for linear systems with good low-rank matrix approximations
by: Lok, Jackie, et al.
Published: (2025)
by: Lok, Jackie, et al.
Published: (2025)
Explicit Construction of Approximate Kolmogorov Superpositions with C2 Smoothness
by: Song, Lunji, et al.
Published: (2025)
by: Song, Lunji, et al.
Published: (2025)
A resource-efficient model for deep kernel learning
by: D'Amore, Luisa
Published: (2024)
by: D'Amore, Luisa
Published: (2024)
Parallel-in-time Multilevel Krylov Methods: A Prototype
by: Erlangga, Yogi A.
Published: (2023)
by: Erlangga, Yogi A.
Published: (2023)
CompressedScaffnew: The First Theoretical Double Acceleration of Communication from Local Training and Compression in Distributed Optimization
by: Condat, Laurent, et al.
Published: (2022)
by: Condat, Laurent, et al.
Published: (2022)
Flexible Multi-Dimensional FFTs for Plane Wave Density Functional Theory Codes
by: Popovici, Doru Thom, et al.
Published: (2024)
by: Popovici, Doru Thom, et al.
Published: (2024)
Compression of Currents and Varifolds
by: Paul, Allen, et al.
Published: (2024)
by: Paul, Allen, et al.
Published: (2024)
Design, Configuration, Implementation, and Performance of a Simple 32 Core Raspberry Pi Cluster
by: Cicirello, Vincent A.
Published: (2017)
by: Cicirello, Vincent A.
Published: (2017)
Robust Tensor CUR Decompositions: Rapid Low-Tucker-Rank Tensor Recovery with Sparse Corruption
by: Cai, HanQin, et al.
Published: (2023)
by: Cai, HanQin, et al.
Published: (2023)
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
An $O(\log N)$ Monte Carlo method for periodic Coulomb systems
by: Gao, Xuanzhao, et al.
Published: (2026)
by: Gao, Xuanzhao, et al.
Published: (2026)
Adaptive finite element methods with optimally preconditioned GMRES guarantee optimal complexity
by: Führer, Thomas, et al.
Published: (2026)
by: Führer, Thomas, et al.
Published: (2026)
RandNet-Parareal: a time-parallel PDE solver using Random Neural Networks
by: Gattiglio, Guglielmo, et al.
Published: (2024)
by: Gattiglio, Guglielmo, et al.
Published: (2024)
Real-time chaotic video encryption based on multithreaded parallel confusion and diffusion
by: Jiang, Dong, et al.
Published: (2023)
by: Jiang, Dong, et al.
Published: (2023)
Distributed Parallel Structure-Aware Presolving for Arrowhead Linear Programs
by: Kempke, Nils-Christian, et al.
Published: (2026)
by: Kempke, Nils-Christian, et al.
Published: (2026)
Anderson acceleration with approximate calculations: applications to scientific computing
by: Pasini, Massimiliano Lupo, et al.
Published: (2022)
by: Pasini, Massimiliano Lupo, et al.
Published: (2022)
Compilation of Generalized Matrix Chains with Symbolic Sizes
by: López, Francisco, et al.
Published: (2025)
by: López, Francisco, et al.
Published: (2025)
Similar Items
-
Matrix-Free Evaluation of High-Order Shifted Boundary Finite Element Operators
by: Wichrowski, Michał
Published: (2025) -
Stochastic trace estimation for parameter-dependent matrices applied to spectral density approximation
by: Matti, Fabio, et al.
Published: (2025) -
Parallelization Strategies for the Randomized Kaczmarz Algorithm on Large-Scale Dense Systems
by: Ferreira, Inês, et al.
Published: (2024) -
Massively Parallel Reductions in Multivariate Polynomial Systems: Bridging the Symbolic Preprocessing Gap on GPGPU Architectures
by: Gokavarapu, Chandrasekhar
Published: (2026) -
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
by: Ansari, Mufakir Qamar, et al.
Published: (2025)