Fast Truncated SVD of Sparse and Dense Matrices on Graphics Processors
Fuente:
arXiv
Saved in:
| Main Authors: | Tomas, Andres E., Quintana-Orti, Enrique S., Anzt, Hartwig |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Alea-BFT: Practical Asynchronous Byzantine Fault Tolerance
by: Antunes, Diogo S., et al.
Published: (2024)
by: Antunes, Diogo S., et al.
Published: (2024)
Characterizing and Fixing Silent Data Loss in Spark-on-AWS-Lambda with Open Table Formats
by: Gandla, Srujan Kumar
Published: (2026)
by: Gandla, Srujan Kumar
Published: (2026)
Mapping Parallel Matrix Multiplication in GotoBLAS2 to the AMD Versal ACAP for Deep Learning
by: Lei, Jie, et al.
Published: (2024)
by: Lei, Jie, et al.
Published: (2024)
Implementation and Evaluation of Fast Raft for Hierarchical Consensus
by: Melnychuk, Anton, et al.
Published: (2025)
by: Melnychuk, Anton, et al.
Published: (2025)
Rebooting Microreboot: Architectural Support for Safe, Parallel Recovery in Microservice Systems
by: Bindschaedler, Laurent
Published: (2026)
by: Bindschaedler, Laurent
Published: (2026)
Investigating Matrix Repartitioning to Address the Over- and Undersubscription Challenge for a GPU-based CFD Solver
by: Olenik, Gregor, et al.
Published: (2025)
by: Olenik, Gregor, et al.
Published: (2025)
Distributed-memory Algorithms for Sparse Matrix Permutation, Extraction, and Assignment
by: Hassani, Elaheh, et al.
Published: (2025)
by: Hassani, Elaheh, et al.
Published: (2025)
DMRlib: Easy-coding and Efficient Resource Management for Job Malleability
by: Iserte, Sergio, et al.
Published: (2026)
by: Iserte, Sergio, et al.
Published: (2026)
From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs
by: Kang, Daemyung, et al.
Published: (2026)
by: Kang, Daemyung, et al.
Published: (2026)
Scaling Point-based Differentiable Rendering for Large-scale Reconstruction
by: Zhao, Hexu, et al.
Published: (2025)
by: Zhao, Hexu, et al.
Published: (2025)
LOw-cOst yet High-Performant Sparse Matrix-Matrix Multiplication on Arm SME Architectures
by: Lei, Kelun, et al.
Published: (2025)
by: Lei, Kelun, et al.
Published: (2025)
From Consensus to Chaos: A Vulnerability Assessment of the RAFT Algorithm
by: Afifi, Tamer, et al.
Published: (2026)
by: Afifi, Tamer, et al.
Published: (2026)
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
by: Homola, Jakub, et al.
Published: (2025)
by: Homola, Jakub, et al.
Published: (2025)
Communication Lower Bounds and Algorithms for Sketching with Random Dense Matrices
by: Daas, Hussam Al, et al.
Published: (2026)
by: Daas, Hussam Al, et al.
Published: (2026)
Improving Locality in Sparse and Dense Matrix Multiplications
by: Dezfuli, Mohammad Mahdi Salehi, et al.
Published: (2024)
by: Dezfuli, Mohammad Mahdi Salehi, et al.
Published: (2024)
Exceeding the Numerical and Performance Characteristics of IEEE-754 SGEMM with BFloat16 Tensor Cores on GPUs for Scientific Computing
by: Bayraktar, Harun, et al.
Published: (2026)
by: Bayraktar, Harun, et al.
Published: (2026)
Serinv: A Scalable Library for the Selected Inversion of Block-Tridiagonal with Arrowhead Matrices
by: Maillou, Vincent, et al.
Published: (2025)
by: Maillou, Vincent, et al.
Published: (2025)
Parallel Reduced Order Modeling for Digital Twins using High-Performance Computing Workflows
by: de Parga, S. Ares, et al.
Published: (2024)
by: de Parga, S. Ares, et al.
Published: (2024)
Staging Blocked Evaluation over Structured Sparse Matrices
by: Das, Pratyush, et al.
Published: (2024)
by: Das, Pratyush, et al.
Published: (2024)
GVE-Louvain: Fast Louvain Algorithm for Community Detection in Shared Memory Setting
by: Sahu, Subhajit
Published: (2023)
by: Sahu, Subhajit
Published: (2023)
GVE-Leiden: Fast Leiden Algorithm for Community Detection in Shared Memory Setting
by: Sahu, Subhajit
Published: (2023)
by: Sahu, Subhajit
Published: (2023)
A Simple Communication Scheme for Distributed Fast Multipole Methods
by: Kailasa, Srinath
Published: (2026)
by: Kailasa, Srinath
Published: (2026)
GVE-LPA: Fast Label Propagation Algorithm (LPA) for Community Detection in Shared Memory Setting
by: Sahu, Subhajit
Published: (2023)
by: Sahu, Subhajit
Published: (2023)
Sparse Checkpointing for Fast and Reliable MoE Training
by: Gandhi, Swapnil, et al.
Published: (2024)
by: Gandhi, Swapnil, et al.
Published: (2024)
Morlet wavelet transform using attenuated sliding Fourier transform and kernel integral for graphic processing unit
by: Yamashita, Yukihiko, et al.
Published: (2021)
by: Yamashita, Yukihiko, et al.
Published: (2021)
TOB-SVD: Total-Order Broadcast with Single-Vote Decisions in the Sleepy Model
by: D'Amato, Francesco, et al.
Published: (2023)
by: D'Amato, Francesco, et al.
Published: (2023)
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
by: Abraham, Ojima, et al.
Published: (2026)
by: Abraham, Ojima, et al.
Published: (2026)
On the Power of Graphical Reconfigurable Circuits
by: Emek, Yuval, et al.
Published: (2024)
by: Emek, Yuval, et al.
Published: (2024)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
by: Tuteja, Keshvi, et al.
Published: (2025)
by: Tuteja, Keshvi, et al.
Published: (2025)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
by: Dwaraknath, Rajat Vadiraj, et al.
Published: (2026)
by: Dwaraknath, Rajat Vadiraj, et al.
Published: (2026)
Multi-level Memory-Centric Profiling on ARM Processors with ARM SPE
by: Miksits, Samuel, et al.
Published: (2024)
by: Miksits, Samuel, et al.
Published: (2024)
Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
by: Shi, Ruimin, et al.
Published: (2026)
by: Shi, Ruimin, et al.
Published: (2026)
Practical GPU Choices for Earth Observation: ResNet-50 Training Throughput on Integrated, Laptop, and Cloud Accelerators
by: Chaturvedi, Ritvik
Published: (2025)
by: Chaturvedi, Ritvik
Published: (2025)
Energy-Aware Scheduling Strategies for Partially-Replicable Task Chains on Heterogeneous Processors
by: Idouar, Yacine, et al.
Published: (2025)
by: Idouar, Yacine, et al.
Published: (2025)
High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
by: Shi, Ruimin, et al.
Published: (2026)
by: Shi, Ruimin, et al.
Published: (2026)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
by: He, Jiaao, et al.
Published: (2024)
by: He, Jiaao, et al.
Published: (2024)
ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
RACS-SADL: Robust and Understandable Randomized Consensus in the Cloud
by: Tennage, Pasindu, et al.
Published: (2024)
by: Tennage, Pasindu, et al.
Published: (2024)
Baxos: Backing off for Robust and Efficient Consensus
by: Tennage, Pasindu, et al.
Published: (2022)
by: Tennage, Pasindu, et al.
Published: (2022)
Similar Items
-
Alea-BFT: Practical Asynchronous Byzantine Fault Tolerance
by: Antunes, Diogo S., et al.
Published: (2024) -
Characterizing and Fixing Silent Data Loss in Spark-on-AWS-Lambda with Open Table Formats
by: Gandla, Srujan Kumar
Published: (2026) -
Mapping Parallel Matrix Multiplication in GotoBLAS2 to the AMD Versal ACAP for Deep Learning
by: Lei, Jie, et al.
Published: (2024) -
Implementation and Evaluation of Fast Raft for Hierarchical Consensus
by: Melnychuk, Anton, et al.
Published: (2025) -
Rebooting Microreboot: Architectural Support for Safe, Parallel Recovery in Microservice Systems
by: Bindschaedler, Laurent
Published: (2026)