Solving Large Rank-Deficient Linear Least-Squares Problems on Shared-Memory CPU Architectures and GPU Architectures
Fuente:
arXiv
Saved in:
| Main Authors: | Chillarón, Mónica, Quintana-Ortí, Gregorio, Vidal, Vicente, Martinsson, Per-Gunnar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automated Synthesis of Quantum Algorithms via Classical Numerical Techniques
by: Huang, Yuxin, et al.
Published: (2024)
by: Huang, Yuxin, et al.
Published: (2024)
Fast GPU Linear Algebra via Compile Time Expression Fusion
by: Curtin, Ryan R., et al.
Published: (2026)
by: Curtin, Ryan R., et al.
Published: (2026)
Trilinos: Enabling Scientific Computing Across Diverse Hardware Architectures at Scale
by: Mayr, Matthias, et al.
Published: (2025)
by: Mayr, Matthias, et al.
Published: (2025)
Boxplots and quartile plots for grouped and periodic angular data
by: Berlinski, Joshua D., et al.
Published: (2026)
by: Berlinski, Joshua D., et al.
Published: (2026)
Design, Configuration, Implementation, and Performance of a Simple 32 Core Raspberry Pi Cluster
by: Cicirello, Vincent A.
Published: (2017)
by: Cicirello, Vincent A.
Published: (2017)
Armadillo: An Efficient Framework for Numerical Linear Algebra
by: Sanderson, Conrad, et al.
Published: (2025)
by: Sanderson, Conrad, et al.
Published: (2025)
Direct Low-Dose CT Image Reconstruction on GPU using Out-Of-Core: Precision and Quality Study
by: Chillarón, M., et al.
Published: (2024)
by: Chillarón, M., et al.
Published: (2024)
Bandicoot: A Templated C++ Library for GPU Linear Algebra
by: Curtin, Ryan R., et al.
Published: (2025)
by: Curtin, Ryan R., et al.
Published: (2025)
A Virtual Processor brings back the Free Lunch
by: Kutschbach, Haymo
Published: (2026)
by: Kutschbach, Haymo
Published: (2026)
Algorithms for Generating Small Random Samples
by: Cicirello, Vincent A.
Published: (2024)
by: Cicirello, Vincent A.
Published: (2024)
Fast Gaussian Distributed Pseudorandom Number Generation in Java via the Ziggurat Algorithm
by: Cicirello, Vincent A.
Published: (2024)
by: Cicirello, Vincent A.
Published: (2024)
Canonicalization of Batched Einstein Summations for Tuning Retrieval
by: Kulkarni, Kaushik, et al.
Published: (2026)
by: Kulkarni, Kaushik, et al.
Published: (2026)
On the Average Runtime of an Open Source Binomial Random Variate Generation Algorithm
by: Cicirello, Vincent A.
Published: (2024)
by: Cicirello, Vincent A.
Published: (2024)
A Two-Level Direct Solver for the Hierarchical Poincaré-Steklov Method
by: Kump, Joseph, et al.
Published: (2025)
by: Kump, Joseph, et al.
Published: (2025)
Assembly of FETI dual operator using CUDA
by: Homola, Jakub, et al.
Published: (2025)
by: Homola, Jakub, et al.
Published: (2025)
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
by: Homola, Jakub, et al.
Published: (2025)
by: Homola, Jakub, et al.
Published: (2025)
Mixed Precision FGMRES-Based Iterative Refinement for Weighted Least Squares
by: Carson, Erin, et al.
Published: (2024)
by: Carson, Erin, et al.
Published: (2024)
Local Adjoints for Simultaneous Preaccumulations with Shared Inputs
by: Blühdorn, Johannes, et al.
Published: (2024)
by: Blühdorn, Johannes, et al.
Published: (2024)
Comparison of substructured non-overlapping domain decomposition and overlapping additive Schwarz methods for large-scale Helmholtz problems with multiple sources
by: Martin, Boris, et al.
Published: (2025)
by: Martin, Boris, et al.
Published: (2025)
Spatial Clustering Approach for Vessel Path Identification
by: Abuella, Mohamed, et al.
Published: (2024)
by: Abuella, Mohamed, et al.
Published: (2024)
Precision-Aware Iterative Algorithms Based on Group-Shared Exponents of Floating-Point Numbers
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing
by: Martinez-Ferrer, Pedro J., et al.
Published: (2025)
by: Martinez-Ferrer, Pedro J., et al.
Published: (2025)
The ensmallen library for flexible numerical optimization
by: Curtin, Ryan R., et al.
Published: (2021)
by: Curtin, Ryan R., et al.
Published: (2021)
Efficient Parallel Scheduling for Sparse Triangular Solvers
by: Böhnlein, Toni, et al.
Published: (2025)
by: Böhnlein, Toni, et al.
Published: (2025)
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
Occlusion aware obstacle prediction using people as sensors
by: Ranaraja, Sithija
Published: (2024)
by: Ranaraja, Sithija
Published: (2024)
Interpretable factorization of clinical questionnaires to identify latent factors of psychopathology
by: Lam, Ka Chun, et al.
Published: (2023)
by: Lam, Ka Chun, et al.
Published: (2023)
Faster Algorithms for Structured Matrix Multiplication via Flip Graph Search
by: Khoruzhii, Kirill, et al.
Published: (2025)
by: Khoruzhii, Kirill, et al.
Published: (2025)
Hybrid parallel discrete adjoints in SU2
by: Blühdorn, Johannes, et al.
Published: (2024)
by: Blühdorn, Johannes, et al.
Published: (2024)
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Stencil Computations on AMD and Nvidia Graphics Processors: Performance and Tuning Strategies
by: Pekkilä, Johannes, et al.
Published: (2024)
by: Pekkilä, Johannes, et al.
Published: (2024)
Incremental Hierarchical Tucker Decomposition
by: Aksoy, Doruk, et al.
Published: (2024)
by: Aksoy, Doruk, et al.
Published: (2024)
Optimizing Tensor Train Decomposition in DNNs for RISC-V Architectures Using Design Space Exploration and Compiler Optimizations
by: Anthimopoulos, Theologos, et al.
Published: (2026)
by: Anthimopoulos, Theologos, et al.
Published: (2026)
Fat API bindings of C++ objects into scripting languages
by: Standish, Russell K.
Published: (2024)
by: Standish, Russell K.
Published: (2024)
Flexible Quaternion Generalized Minimal Residual Method for Ill-Posed Quaternion Inverse Problems
by: Liu, Xuan, et al.
Published: (2024)
by: Liu, Xuan, et al.
Published: (2024)
Support data for the paper "Addressing Large Rank-Deficient Linear Least-Squares Problems on Shared-Memory CPU and GPU Architectures"
by: Quintana-Ortí, Gregorio, et al.
Published: (2026)
by: Quintana-Ortí, Gregorio, et al.
Published: (2026)
Which graph motif parameters count?
by: Bläser, Markus, et al.
Published: (2025)
by: Bläser, Markus, et al.
Published: (2025)
A Methodological Analysis of Empirical Studies in Quantum Software Testing
by: Li, Yuechen, et al.
Published: (2026)
by: Li, Yuechen, et al.
Published: (2026)
Fast Evaluation of Truncated Neumann Series by Low-Product Radix Kernels
by: Sao, Piyush
Published: (2026)
by: Sao, Piyush
Published: (2026)
Similar Items
-
Automated Synthesis of Quantum Algorithms via Classical Numerical Techniques
by: Huang, Yuxin, et al.
Published: (2024) -
Fast GPU Linear Algebra via Compile Time Expression Fusion
by: Curtin, Ryan R., et al.
Published: (2026) -
Trilinos: Enabling Scientific Computing Across Diverse Hardware Architectures at Scale
by: Mayr, Matthias, et al.
Published: (2025) -
Boxplots and quartile plots for grouped and periodic angular data
by: Berlinski, Joshua D., et al.
Published: (2026) -
Design, Configuration, Implementation, and Performance of a Simple 32 Core Raspberry Pi Cluster
by: Cicirello, Vincent A.
Published: (2017)