Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
Fuente:
arXiv
Guardado en:
| Autores principales: | Ringoot, Evelyne, Alomairy, Rabab, Edelman, Alan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
por: Ringoot, Evelyne, et al.
Publicado: (2025)
por: Ringoot, Evelyne, et al.
Publicado: (2025)
Hierarchical Recursive Precision for Accelerating Symmetric Linear Solves on MXUs
por: Carrica, Vicki, et al.
Publicado: (2026)
por: Carrica, Vicki, et al.
Publicado: (2026)
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
por: Carrica, Vicki, et al.
Publicado: (2025)
por: Carrica, Vicki, et al.
Publicado: (2025)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
por: Panova, Elena, et al.
Publicado: (2022)
por: Panova, Elena, et al.
Publicado: (2022)
Integrating Odeint Time Stepping into OpenFPM for Distributed and GPU Accelerated Numerical Solvers
por: Singh, Abhinav, et al.
Publicado: (2023)
por: Singh, Abhinav, et al.
Publicado: (2023)
A Communication Avoiding and Reducing Algorithm for Symmetric Eigenproblem for Very Small Matrices
por: Katagiri, Takahiro, et al.
Publicado: (2024)
por: Katagiri, Takahiro, et al.
Publicado: (2024)
Leveraging Hardware-Aware Computation in Mixed-Precision Matrix Multiply: A Tile-Centric Approach
por: Zhang, Qiao, et al.
Publicado: (2025)
por: Zhang, Qiao, et al.
Publicado: (2025)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
por: Tu, Jiqun, et al.
Publicado: (2026)
por: Tu, Jiqun, et al.
Publicado: (2026)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
por: Bergach, Mohamed Amine
Publicado: (2026)
por: Bergach, Mohamed Amine
Publicado: (2026)
Enabling mixed-precision in spectral element codes
por: Chen, Yanxiang, et al.
Publicado: (2025)
por: Chen, Yanxiang, et al.
Publicado: (2025)
Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
por: Wang, Hansheng, et al.
Publicado: (2025)
por: Wang, Hansheng, et al.
Publicado: (2025)
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
por: Villalobos, Johansell, et al.
Publicado: (2025)
por: Villalobos, Johansell, et al.
Publicado: (2025)
On the Challenges of Energy-Efficiency Analysis in HPC Systems: Evaluating Synthetic Benchmarks and Gromacs
por: Machado, Rafael Ravedutti Lucio, et al.
Publicado: (2025)
por: Machado, Rafael Ravedutti Lucio, et al.
Publicado: (2025)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
por: Li, Yifan, et al.
Publicado: (2026)
por: Li, Yifan, et al.
Publicado: (2026)
High-Performance Star-M SVD for Big Data Compression
por: Hussain, Md Taufique, et al.
Publicado: (2026)
por: Hussain, Md Taufique, et al.
Publicado: (2026)
A shared compilation stack for distributed-memory parallelism in stencil DSLs
por: Bisbas, George, et al.
Publicado: (2024)
por: Bisbas, George, et al.
Publicado: (2024)
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
por: Zhu, Honglin, et al.
Publicado: (2026)
por: Zhu, Honglin, et al.
Publicado: (2026)
Robustness and Accuracy in Pipelined Bi-Conjugate Gradient Stabilized Method: A Comparative Study
por: Havdiak, Mykhailo, et al.
Publicado: (2024)
por: Havdiak, Mykhailo, et al.
Publicado: (2024)
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
por: Bellavita, Julian, et al.
Publicado: (2026)
por: Bellavita, Julian, et al.
Publicado: (2026)
Efficient N-to-M Checkpointing Algorithm for Finite Element Simulations
por: Ham, David A., et al.
Publicado: (2024)
por: Ham, David A., et al.
Publicado: (2024)
A new open source framework for multiscale modeling of fibrous materials on heterogeneous supercomputers
por: Merson, Jacob, et al.
Publicado: (2023)
por: Merson, Jacob, et al.
Publicado: (2023)
Enabling MPI communication within Numba/LLVM JIT-compiled Python code using numba-mpi v1.0
por: Derlatka, Kacper, et al.
Publicado: (2024)
por: Derlatka, Kacper, et al.
Publicado: (2024)
SYCL compute kernels for ExaHyPE
por: Loi, Chung Ming, et al.
Publicado: (2023)
por: Loi, Chung Ming, et al.
Publicado: (2023)
Evaluation of POSIT Arithmetic with Accelerators
por: Nakasato, Naohito, et al.
Publicado: (2024)
por: Nakasato, Naohito, et al.
Publicado: (2024)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
por: Li, Junjie
Publicado: (2024)
por: Li, Junjie
Publicado: (2024)
A Graph-Based, Distributed Memory, Modeling Abstraction for Optimization
por: Cole, David L., et al.
Publicado: (2025)
por: Cole, David L., et al.
Publicado: (2025)
Binsparse: A Specification for Cross-Platform Storage of Sparse Matrices and Tensors
por: Brock, Benjamin, et al.
Publicado: (2025)
por: Brock, Benjamin, et al.
Publicado: (2025)
Enabling mixed-precision with the help of tools: A Nekbone case study
por: Chen, Yanxiang, et al.
Publicado: (2024)
por: Chen, Yanxiang, et al.
Publicado: (2024)
On the energy efficiency of sparse matrix computations on multi-GPU clusters
por: Bernaschi, Massimo, et al.
Publicado: (2025)
por: Bernaschi, Massimo, et al.
Publicado: (2025)
NApy: Efficient Statistics in Python for Large-Scale Heterogeneous Data with Enhanced Support for Missing Data
por: Woller, Fabian, et al.
Publicado: (2025)
por: Woller, Fabian, et al.
Publicado: (2025)
Performance measurements of modern Fortran MPI applications with Score-P
por: Corbin, Gregor
Publicado: (2025)
por: Corbin, Gregor
Publicado: (2025)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
por: Katagiri, Takahiro, et al.
Publicado: (2024)
por: Katagiri, Takahiro, et al.
Publicado: (2024)
Automated MPI-X code generation for scalable finite-difference solvers
por: Bisbas, George, et al.
Publicado: (2023)
por: Bisbas, George, et al.
Publicado: (2023)
Secure and Parallel Determinant Computation for Large-Scale Matrices in Edge Environments
por: Panth, Prajwal
Publicado: (2026)
por: Panth, Prajwal
Publicado: (2026)
GoldbachGPU: An Open Source GPU-Accelerated Framework for Verification of Goldbach's Conjecture
por: Llorente-Saguer, Isaac
Publicado: (2026)
por: Llorente-Saguer, Isaac
Publicado: (2026)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
por: Tuteja, Keshvi, et al.
Publicado: (2025)
por: Tuteja, Keshvi, et al.
Publicado: (2025)
Unlocking massively parallel spectral proper orthogonal decompositions in the PySPOD package
por: Rogowski, Marcin, et al.
Publicado: (2023)
por: Rogowski, Marcin, et al.
Publicado: (2023)
PennyLane-Lightning MPI: A massively scalable quantum circuit simulator based on distributed computing in CPU clusters
por: Kang, Ji-Hoon, et al.
Publicado: (2025)
por: Kang, Ji-Hoon, et al.
Publicado: (2025)
Performance Evaluation of General Purpose Large Language Models for Basic Linear Algebra Subprograms Code Generation
por: Mukunoki, Daichi, et al.
Publicado: (2025)
por: Mukunoki, Daichi, et al.
Publicado: (2025)
SPUMA: a minimally invasive approach to the GPU porting of OPENFOAM
por: Bnà, Simone, et al.
Publicado: (2025)
por: Bnà, Simone, et al.
Publicado: (2025)
Ejemplares similares
-
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
por: Ringoot, Evelyne, et al.
Publicado: (2025) -
Hierarchical Recursive Precision for Accelerating Symmetric Linear Solves on MXUs
por: Carrica, Vicki, et al.
Publicado: (2026) -
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
por: Carrica, Vicki, et al.
Publicado: (2025) -
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
por: Panova, Elena, et al.
Publicado: (2022) -
Integrating Odeint Time Stepping into OpenFPM for Distributed and GPU Accelerated Numerical Solvers
por: Singh, Abhinav, et al.
Publicado: (2023)