AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Chendi, Jia, Haipeng, Cao, Hang, Yao, Jianyu, Shi, Boqian, Xiang, Chunyang, Sun, Jinbo, Lu, Pengqi, Zhang, Yunquan
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929684194787328
author Li, Chendi
Jia, Haipeng
Cao, Hang
Yao, Jianyu
Shi, Boqian
Xiang, Chunyang
Sun, Jinbo
Lu, Pengqi
Zhang, Yunquan
author_facet Li, Chendi
Jia, Haipeng
Cao, Hang
Yao, Jianyu
Shi, Boqian
Xiang, Chunyang
Sun, Jinbo
Lu, Pengqi
Zhang, Yunquan
contents In recent years, general matrix-matrix multiplication with non-regular-shaped input matrices has been widely used in many applications like deep learning and has drawn more and more attention. However, conventional implementations are not suited for non-regular-shaped matrix-matrix multiplications, and few works focus on optimizing tall-and-skinny matrix-matrix multiplication on CPUs. This paper proposes an auto-tuning framework, AutoTSMM, to build high-performance tall-and-skinny matrix-matrix multiplication. AutoTSMM selects the optimal inner kernels in the install-time stage and generates an execution plan for the pre-pack tall-and-skinny matrix-matrix multiplication in the runtime stage. Experiments demonstrate that AutoTSMM achieves competitive performance comparing to state-of-the-art tall-and-skinny matrix-matrix multiplication. And, it outperforms all conventional matrix-matrix multiplication implementations.
format Preprint
id arxiv_https___arxiv_org_abs_2208_08088
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
Li, Chendi
Jia, Haipeng
Cao, Hang
Yao, Jianyu
Shi, Boqian
Xiang, Chunyang
Sun, Jinbo
Lu, Pengqi
Zhang, Yunquan
Distributed, Parallel, and Cluster Computing
D.1.3
In recent years, general matrix-matrix multiplication with non-regular-shaped input matrices has been widely used in many applications like deep learning and has drawn more and more attention. However, conventional implementations are not suited for non-regular-shaped matrix-matrix multiplications, and few works focus on optimizing tall-and-skinny matrix-matrix multiplication on CPUs. This paper proposes an auto-tuning framework, AutoTSMM, to build high-performance tall-and-skinny matrix-matrix multiplication. AutoTSMM selects the optimal inner kernels in the install-time stage and generates an execution plan for the pre-pack tall-and-skinny matrix-matrix multiplication in the runtime stage. Experiments demonstrate that AutoTSMM achieves competitive performance comparing to state-of-the-art tall-and-skinny matrix-matrix multiplication. And, it outperforms all conventional matrix-matrix multiplication implementations.
title AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
topic Distributed, Parallel, and Cluster Computing
D.1.3
url https://arxiv.org/abs/2208.08088