AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866929684194787328 |
|---|---|
| author | Li, Chendi Jia, Haipeng Cao, Hang Yao, Jianyu Shi, Boqian Xiang, Chunyang Sun, Jinbo Lu, Pengqi Zhang, Yunquan |
| author_facet | Li, Chendi Jia, Haipeng Cao, Hang Yao, Jianyu Shi, Boqian Xiang, Chunyang Sun, Jinbo Lu, Pengqi Zhang, Yunquan |
| contents | In recent years, general matrix-matrix multiplication with non-regular-shaped input matrices has been widely used in many applications like deep learning and has drawn more and more attention. However, conventional implementations are not suited for non-regular-shaped matrix-matrix multiplications, and few works focus on optimizing tall-and-skinny matrix-matrix multiplication on CPUs. This paper proposes an auto-tuning framework, AutoTSMM, to build high-performance tall-and-skinny matrix-matrix multiplication. AutoTSMM selects the optimal inner kernels in the install-time stage and generates an execution plan for the pre-pack tall-and-skinny matrix-matrix multiplication in the runtime stage. Experiments demonstrate that AutoTSMM achieves competitive performance comparing to state-of-the-art tall-and-skinny matrix-matrix multiplication. And, it outperforms all conventional matrix-matrix multiplication implementations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2208_08088 |
| institution | arXiv |
| publishDate | 2022 |
| record_format | arxiv |
| spellingShingle | AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs Li, Chendi Jia, Haipeng Cao, Hang Yao, Jianyu Shi, Boqian Xiang, Chunyang Sun, Jinbo Lu, Pengqi Zhang, Yunquan Distributed, Parallel, and Cluster Computing D.1.3 In recent years, general matrix-matrix multiplication with non-regular-shaped input matrices has been widely used in many applications like deep learning and has drawn more and more attention. However, conventional implementations are not suited for non-regular-shaped matrix-matrix multiplications, and few works focus on optimizing tall-and-skinny matrix-matrix multiplication on CPUs. This paper proposes an auto-tuning framework, AutoTSMM, to build high-performance tall-and-skinny matrix-matrix multiplication. AutoTSMM selects the optimal inner kernels in the install-time stage and generates an execution plan for the pre-pack tall-and-skinny matrix-matrix multiplication in the runtime stage. Experiments demonstrate that AutoTSMM achieves competitive performance comparing to state-of-the-art tall-and-skinny matrix-matrix multiplication. And, it outperforms all conventional matrix-matrix multiplication implementations. |
| title | AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs |
| topic | Distributed, Parallel, and Cluster Computing D.1.3 |
| url | https://arxiv.org/abs/2208.08088 |