FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhu, Honglin, Cao, Jiaping, Shao, Jiang, Feng, Siyuan, Qiu, Qian, Chen, Peng, Zhang, Xu, Zhou, Yixian, Yiu, Man Lung, Ji, Guang, Deng, Minwen, Zhu, Wenxi, Meng, Jintao
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910210785804288
author Zhu, Honglin
Cao, Jiaping
Shao, Jiang
Feng, Siyuan
Qiu, Qian
Chen, Peng
Zhang, Xu
Zhou, Yixian
Yiu, Man Lung
Ji, Guang
Deng, Minwen
Zhu, Wenxi
Meng, Jintao
author_facet Zhu, Honglin
Cao, Jiaping
Shao, Jiang
Feng, Siyuan
Qiu, Qian
Chen, Peng
Zhang, Xu
Zhou, Yixian
Yiu, Man Lung
Ji, Guang
Deng, Minwen
Zhu, Wenxi
Meng, Jintao
contents Peak breaking Matrix Multiplication is a promising technique to improve the performance of DL, especially in LLM training and inference. We present FalconGEMM, a cross-platform framework that automates the deployment, optimization, and selection of Lower-Complexity Matrix Multiplication Algorithms (LCMAs) across diverse hardware. There are three key innovations: (1) a Deployment Module that enables portable execution across various hardware and input configurations through code generation; (2) an Execution Module with Group-Parallel Optimizations that maximizes on-chip data reuse, utilizes parallel resources, and reduces bandwidth overhead; and (3) a Decision Module featuring a lightweight analytical performance model to select the optimal strategy based on matrix shapes and hardware profiles. Extensive evaluation is conducted on LLM workloads across GPU (H20, A100) and CPU (ARM, x86) architectures with multiple data types. FalconGEMM succeeds in delivering peak breaking performance and outperforms GEMM libraries (e.g., cuBLAS, CUTLASS, Intel MKL, etc) by 7.59%-17.85% and LCMA competitors like AlphaTensor by 12.41%-55.61%. Our framework makes the theoretical promise of LCMAs practical for production deployment across the heterogeneous landscape of modern hardware.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06057
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
Zhu, Honglin
Cao, Jiaping
Shao, Jiang
Feng, Siyuan
Qiu, Qian
Chen, Peng
Zhang, Xu
Zhou, Yixian
Yiu, Man Lung
Ji, Guang
Deng, Minwen
Zhu, Wenxi
Meng, Jintao
Distributed, Parallel, and Cluster Computing
Mathematical Software
Peak breaking Matrix Multiplication is a promising technique to improve the performance of DL, especially in LLM training and inference. We present FalconGEMM, a cross-platform framework that automates the deployment, optimization, and selection of Lower-Complexity Matrix Multiplication Algorithms (LCMAs) across diverse hardware. There are three key innovations: (1) a Deployment Module that enables portable execution across various hardware and input configurations through code generation; (2) an Execution Module with Group-Parallel Optimizations that maximizes on-chip data reuse, utilizes parallel resources, and reduces bandwidth overhead; and (3) a Decision Module featuring a lightweight analytical performance model to select the optimal strategy based on matrix shapes and hardware profiles. Extensive evaluation is conducted on LLM workloads across GPU (H20, A100) and CPU (ARM, x86) architectures with multiple data types. FalconGEMM succeeds in delivering peak breaking performance and outperforms GEMM libraries (e.g., cuBLAS, CUTLASS, Intel MKL, etc) by 7.59%-17.85% and LCMA competitors like AlphaTensor by 12.41%-55.61%. Our framework makes the theoretical promise of LCMAs practical for production deployment across the heterogeneous landscape of modern hardware.
title FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
topic Distributed, Parallel, and Cluster Computing
Mathematical Software
url https://arxiv.org/abs/2605.06057