| _version_ | 1866902216664678400 |
|---|---|
| author | Wang, Hansheng |
| author_facet | Wang, Hansheng |
| contents | <p>The two-stage eigenvalue decomposition (EVD) method outperforms conventional one-stage method on GPUs and heterogeneous architectures, especially when eigenvectors are not required. However, its performance advantage diminishes when performing back transformation to obtain eigenvectors. To address this, we propose two key solutions: 1) replacing BLAS3 operations with BLAS2 operations during the bulge-chasing back transformation for better performance, and 2) reordering the back transformation workflow from a backward pattern to a new parallelism-driven pattern to hide divide-and-conquer latency, at the cost of one additional GEMM computation.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_17075126 |
| institution | Zenodo |
| language | |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | RBT In Two-stage EVD Wang, Hansheng <p>The two-stage eigenvalue decomposition (EVD) method outperforms conventional one-stage method on GPUs and heterogeneous architectures, especially when eigenvectors are not required. However, its performance advantage diminishes when performing back transformation to obtain eigenvectors. To address this, we propose two key solutions: 1) replacing BLAS3 operations with BLAS2 operations during the bulge-chasing back transformation for better performance, and 2) reordering the back transformation workflow from a backward pattern to a new parallelism-driven pattern to hide divide-and-conquer latency, at the cost of one additional GEMM computation.</p> |
| title | RBT In Two-stage EVD |
| url | https://doi.org/10.5281/zenodo.17075126 |