Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Hansheng, Zhan, Ruiyi, Huang, Dajun, Liu, Xingchen, Li, Qiao, Duan, Hancong, Tao, Dingwen, Tan, Guangming, Zhang, Shaoshuai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917095720091648
author Wang, Hansheng
Zhan, Ruiyi
Huang, Dajun
Liu, Xingchen
Li, Qiao
Duan, Hancong
Tao, Dingwen
Tan, Guangming
Zhang, Shaoshuai
author_facet Wang, Hansheng
Zhan, Ruiyi
Huang, Dajun
Liu, Xingchen
Li, Qiao
Duan, Hancong
Tao, Dingwen
Tan, Guangming
Zhang, Shaoshuai
contents Large symmetric eigenvalue problems are commonly observed in many disciplines such as Chemistry and Physics, and several libraries including cuSOLVERMp, MAGMA and ELPA support computing large eigenvalue decomposition on multi-GPU or multi-CPU-GPU hybrid architectures. However, these libraries do not provide satisfied performance that all of the libraries only utilize around 1.5\% of the peak multi-GPU performance. In this paper, we propose a pipelined two-stage eigenvalue decomposition algorithm instead of conventional subsequent algorithm with substantial optimizations. On an 8$\times$A100 platform, our implementation surpasses state-of-the-art cuSOLVERMp and MAGMA baselines, delivering mean speedups of 5.74$\times$ and 6.59$\times$, with better strong and weak scalability.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16174
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
Wang, Hansheng
Zhan, Ruiyi
Huang, Dajun
Liu, Xingchen
Li, Qiao
Duan, Hancong
Tao, Dingwen
Tan, Guangming
Zhang, Shaoshuai
Mathematical Software
Distributed, Parallel, and Cluster Computing
Large symmetric eigenvalue problems are commonly observed in many disciplines such as Chemistry and Physics, and several libraries including cuSOLVERMp, MAGMA and ELPA support computing large eigenvalue decomposition on multi-GPU or multi-CPU-GPU hybrid architectures. However, these libraries do not provide satisfied performance that all of the libraries only utilize around 1.5\% of the peak multi-GPU performance. In this paper, we propose a pipelined two-stage eigenvalue decomposition algorithm instead of conventional subsequent algorithm with substantial optimizations. On an 8$\times$A100 platform, our implementation surpasses state-of-the-art cuSOLVERMp and MAGMA baselines, delivering mean speedups of 5.74$\times$ and 6.59$\times$, with better strong and weak scalability.
title Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
topic Mathematical Software
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2511.16174