Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Shifang, Li, Huiyuan, Sheng, Hongjiao, Gui, Haoyuan, Zhang, Xiaoyu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908491058249728
author Liu, Shifang
Li, Huiyuan
Sheng, Hongjiao
Gui, Haoyuan
Zhang, Xiaoyu
author_facet Liu, Shifang
Li, Huiyuan
Sheng, Hongjiao
Gui, Haoyuan
Zhang, Xiaoyu
contents Singular Value Decomposition (SVD) is a fundamental matrix factorization technique in linear algebra, widely applied in numerous matrix-related problems. However, traditional SVD approaches are hindered by slow panel factorization and frequent CPU-GPU data transfers in heterogeneous systems, despite advancements in GPU computational capabilities. In this paper, we introduce a GPU-centered SVD algorithm, incorporating a novel GPU-based bidiagonal divide-and-conquer (BDC) method. We reformulate the algorithm and data layout of different steps for SVD computation, performing all panel-level computations and trailing matrix updates entirely on GPU to eliminate CPU-GPU data transfers. Furthermore, we integrate related computations to optimize BLAS utilization, thereby increasing arithmetic intensity and fully leveraging the computational capabilities of GPUs. Additionally, we introduce a newly developed GPU-based BDC algorithm that restructures the workflow to eliminate matrix-level CPU-GPU data transfers and enable asynchronous execution between the CPU and GPU. Experimental results on AMD MI210 and NVIDIA V100 GPUs demonstrate that our proposed method achieves speedups of up to 1293.64x/7.47x and 14.10x/12.38x compared to rocSOLVER/cuSOLVER and MAGMA, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11467
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
Liu, Shifang
Li, Huiyuan
Sheng, Hongjiao
Gui, Haoyuan
Zhang, Xiaoyu
Distributed, Parallel, and Cluster Computing
Performance
Singular Value Decomposition (SVD) is a fundamental matrix factorization technique in linear algebra, widely applied in numerous matrix-related problems. However, traditional SVD approaches are hindered by slow panel factorization and frequent CPU-GPU data transfers in heterogeneous systems, despite advancements in GPU computational capabilities. In this paper, we introduce a GPU-centered SVD algorithm, incorporating a novel GPU-based bidiagonal divide-and-conquer (BDC) method. We reformulate the algorithm and data layout of different steps for SVD computation, performing all panel-level computations and trailing matrix updates entirely on GPU to eliminate CPU-GPU data transfers. Furthermore, we integrate related computations to optimize BLAS utilization, thereby increasing arithmetic intensity and fully leveraging the computational capabilities of GPUs. Additionally, we introduce a newly developed GPU-based BDC algorithm that restructures the workflow to eliminate matrix-level CPU-GPU data transfers and enable asynchronous execution between the CPU and GPU. Experimental results on AMD MI210 and NVIDIA V100 GPUs demonstrate that our proposed method achieves speedups of up to 1293.64x/7.47x and 14.10x/12.38x compared to rocSOLVER/cuSOLVER and MAGMA, respectively.
title Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
topic Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2508.11467