GPU-Accelerated Parallel Selected Inversion for Structured Matrices Using sTiles
Fuente:
arXiv
Saved in:
| Main Authors: | Fattah, Esmail Abdul, Ltaief, Hatem, Rue, Havard, Keyes, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
sTiles: An Accelerated Computational Framework for Sparse Factorizations of Structured Matrices
by: Fattah, Esmail Abdul, et al.
Published: (2025)
by: Fattah, Esmail Abdul, et al.
Published: (2025)
PyINLA: Fast Bayesian Inference for Latent Gaussian Models in Python
by: Fattah, Esmail Abdul, et al.
Published: (2026)
by: Fattah, Esmail Abdul, et al.
Published: (2026)
Accelerating AI Performance using Anderson Extrapolation on GPUs
by: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Published: (2024)
by: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Published: (2024)
Parallel Selected Inversion for Space-Time Gaussian Markov Random Fields
by: Zhumekenov, Abylay, et al.
Published: (2023)
by: Zhumekenov, Abylay, et al.
Published: (2023)
GPU-Accelerated Modified Bessel Function of the Second Kind for Gaussian Processes
by: Geng, Zipei, et al.
Published: (2025)
by: Geng, Zipei, et al.
Published: (2025)
GPU-Accelerated Vecchia Approximations of Gaussian Processes for Geospatial Data using Batched Matrix Computations
by: Pan, Qilong, et al.
Published: (2024)
by: Pan, Qilong, et al.
Published: (2024)
Toward Capturing Genetic Epistasis From Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression
by: Ltaief, Hatem, et al.
Published: (2024)
by: Ltaief, Hatem, et al.
Published: (2024)
Accelerating Mixed-Precision Out-of-Core Cholesky Factorization with Static Task Scheduling
by: Ren, Jie, et al.
Published: (2024)
by: Ren, Jie, et al.
Published: (2024)
GigaAPI for GPU Parallelization
by: Suvarna, M., et al.
Published: (2025)
by: Suvarna, M., et al.
Published: (2025)
Parallelizing a modern GPU simulator
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
A Precision Emulation Approach to the GPU Acceleration of Ab Initio Electronic Structure Calculations
by: Liu, Hang, et al.
Published: (2026)
by: Liu, Hang, et al.
Published: (2026)
In-Situ Techniques on GPU-Accelerated Data-Intensive Applications
by: Ju, Yi, et al.
Published: (2024)
by: Ju, Yi, et al.
Published: (2024)
Serinv: A Scalable Library for the Selected Inversion of Block-Tridiagonal with Arrowhead Matrices
by: Maillou, Vincent, et al.
Published: (2025)
by: Maillou, Vincent, et al.
Published: (2025)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
by: Boudaoud, Afif, et al.
Published: (2026)
by: Boudaoud, Afif, et al.
Published: (2026)
A Tale of Three Location Trackers: AirTag, SmartTag, and Tile
by: Jang, HyunSeok Daniel, et al.
Published: (2025)
by: Jang, HyunSeok Daniel, et al.
Published: (2025)
Staging Blocked Evaluation over Structured Sparse Matrices
by: Das, Pratyush, et al.
Published: (2024)
by: Das, Pratyush, et al.
Published: (2024)
ZO2: Scalable Zeroth-Order Fine-Tuning for Extremely Large Language Models with Limited GPU Memory
by: Wang, Liangyu, et al.
Published: (2025)
by: Wang, Liangyu, et al.
Published: (2025)
Disaggregated Design for GPU-Based Volumetric Data Structures
by: Meneghin, Massimiliano, et al.
Published: (2025)
by: Meneghin, Massimiliano, et al.
Published: (2025)
GoldbachGPU: An Open Source GPU-Accelerated Framework for Verification of Goldbach's Conjecture
by: Llorente-Saguer, Isaac
Published: (2026)
by: Llorente-Saguer, Isaac
Published: (2026)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
by: Davis, Joshua H., et al.
Published: (2026)
by: Davis, Joshua H., et al.
Published: (2026)
Parallel Quadratic Selected Inversion in Quantum Transport Simulation
by: Maillou, Vincent, et al.
Published: (2026)
by: Maillou, Vincent, et al.
Published: (2026)
Accelerating Particle-in-Cell Monte Carlo Simulations with MPI, OpenMP/OpenACC and Asynchronous Multi-GPU Programming
by: Williams, Jeremy J., et al.
Published: (2024)
by: Williams, Jeremy J., et al.
Published: (2024)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
by: Israel, Daniel, et al.
Published: (2025)
by: Israel, Daniel, et al.
Published: (2025)
Can Asymmetric Tile Buffering Be Beneficial?
by: Wang, Chengyue, et al.
Published: (2025)
by: Wang, Chengyue, et al.
Published: (2025)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
by: Nicusan, Andrei-Leonard, et al.
Published: (2025)
by: Nicusan, Andrei-Leonard, et al.
Published: (2025)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
by: Li, Zhuojin, et al.
Published: (2025)
by: Li, Zhuojin, et al.
Published: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
by: Liu, Shifang, et al.
Published: (2025)
by: Liu, Shifang, et al.
Published: (2025)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
by: Curless, Brian, et al.
Published: (2025)
by: Curless, Brian, et al.
Published: (2025)
cuTeSpMM: Accelerating Sparse-Dense Matrix Multiplication using GPU Tensor Cores
by: Xiang, Lizhi, et al.
Published: (2025)
by: Xiang, Lizhi, et al.
Published: (2025)
Parallel Approximations for High-Dimensional Multivariate Normal Probability Computation in Confidence Region Detection Applications
by: Zhang, Xiran, et al.
Published: (2024)
by: Zhang, Xiran, et al.
Published: (2024)
GPU Acceleration of Sparse Fully Homomorphic Encrypted DNNs
by: D'Agata, Lara, et al.
Published: (2026)
by: D'Agata, Lara, et al.
Published: (2026)
Scalable GPU Performance Variability Analysis framework
by: Lahiry, Ankur, et al.
Published: (2025)
by: Lahiry, Ankur, et al.
Published: (2025)
On the Partitioning of GPU Power among Multi-Instances
by: Vamja, Tirth, et al.
Published: (2025)
by: Vamja, Tirth, et al.
Published: (2025)
The Energy Cost of Execution-Idle in GPU Clusters
by: Lei, Yiran, et al.
Published: (2026)
by: Lei, Yiran, et al.
Published: (2026)
Binary Bleed: Fast Distributed and Parallel Method for Automatic Model Selection
by: Barron, Ryan, et al.
Published: (2024)
by: Barron, Ryan, et al.
Published: (2024)
gDist: Efficient Distance Computation between 3D Meshes on GPU
by: Fang, Peng, et al.
Published: (2024)
by: Fang, Peng, et al.
Published: (2024)
Taking GPU Programming Models to Task for Performance Portability
by: Davis, Joshua H., et al.
Published: (2024)
by: Davis, Joshua H., et al.
Published: (2024)
The Tile: A 2D Map of Ranking Scores for Two-Class Classification
by: Piérard, Sébastien, et al.
Published: (2024)
by: Piérard, Sébastien, et al.
Published: (2024)
Profiling and optimization of multi-card GPU machine learning jobs
by: Lawenda, Marcin, et al.
Published: (2025)
by: Lawenda, Marcin, et al.
Published: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
by: Zhao, Yanbo, et al.
Published: (2025)
by: Zhao, Yanbo, et al.
Published: (2025)
Similar Items
-
sTiles: An Accelerated Computational Framework for Sparse Factorizations of Structured Matrices
by: Fattah, Esmail Abdul, et al.
Published: (2025) -
PyINLA: Fast Bayesian Inference for Latent Gaussian Models in Python
by: Fattah, Esmail Abdul, et al.
Published: (2026) -
Accelerating AI Performance using Anderson Extrapolation on GPUs
by: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Published: (2024) -
Parallel Selected Inversion for Space-Time Gaussian Markov Random Fields
by: Zhumekenov, Abylay, et al.
Published: (2023) -
GPU-Accelerated Modified Bessel Function of the Second Kind for Gaussian Processes
by: Geng, Zipei, et al.
Published: (2025)