Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917095720091648 |
|---|---|
| author | Wang, Hansheng Zhan, Ruiyi Huang, Dajun Liu, Xingchen Li, Qiao Duan, Hancong Tao, Dingwen Tan, Guangming Zhang, Shaoshuai |
| author_facet | Wang, Hansheng Zhan, Ruiyi Huang, Dajun Liu, Xingchen Li, Qiao Duan, Hancong Tao, Dingwen Tan, Guangming Zhang, Shaoshuai |
| contents | Large symmetric eigenvalue problems are commonly observed in many disciplines such as Chemistry and Physics, and several libraries including cuSOLVERMp, MAGMA and ELPA support computing large eigenvalue decomposition on multi-GPU or multi-CPU-GPU hybrid architectures. However, these libraries do not provide satisfied performance that all of the libraries only utilize around 1.5\% of the peak multi-GPU performance. In this paper, we propose a pipelined two-stage eigenvalue decomposition algorithm instead of conventional subsequent algorithm with substantial optimizations. On an 8$\times$A100 platform, our implementation surpasses state-of-the-art cuSOLVERMp and MAGMA baselines, delivering mean speedups of 5.74$\times$ and 6.59$\times$, with better strong and weak scalability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_16174 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures Wang, Hansheng Zhan, Ruiyi Huang, Dajun Liu, Xingchen Li, Qiao Duan, Hancong Tao, Dingwen Tan, Guangming Zhang, Shaoshuai Mathematical Software Distributed, Parallel, and Cluster Computing Large symmetric eigenvalue problems are commonly observed in many disciplines such as Chemistry and Physics, and several libraries including cuSOLVERMp, MAGMA and ELPA support computing large eigenvalue decomposition on multi-GPU or multi-CPU-GPU hybrid architectures. However, these libraries do not provide satisfied performance that all of the libraries only utilize around 1.5\% of the peak multi-GPU performance. In this paper, we propose a pipelined two-stage eigenvalue decomposition algorithm instead of conventional subsequent algorithm with substantial optimizations. On an 8$\times$A100 platform, our implementation surpasses state-of-the-art cuSOLVERMp and MAGMA baselines, delivering mean speedups of 5.74$\times$ and 6.59$\times$, with better strong and weak scalability. |
| title | Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures |
| topic | Mathematical Software Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2511.16174 |