MAC Performance and Algorithmic Optimization in Matrix Multiplication Workloads

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Negasa, D, Mohamed, N. O, Ghani, Arfan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910277377720320
author Negasa, D
Mohamed, N. O
Ghani, Arfan
author_facet Negasa, D
Mohamed, N. O
Ghani, Arfan
contents Matrix multiplication is a fundamental computational kernel underlying a wide range of real-world applications, including machine learning, scientific computing, signal processing, and computer graphics. Its performance directly impacts the efficiency, scalability, and energy consumption of modern computing systems. This paper presents a comparative analysis of several matrix multiplication algorithms implemented in software and examined in the context of their hardware execution characteristics. Naive, NumPy, Strassen, and Winograd algorithms are evaluated based on execution time, user time, and CPU time across increasing matrix sizes. The performance metrics reveal computational bottlenecks and highlight the benefits of algorithmic optimizations. Furthermore, the study investigates the mathematical operations underlying each algorithm and analyzes how matrix dimensions influence MAC (Multiply-Accumulate) behavior and overall computational efficiency in the hardware domain. The results provide a performance benchmark and contribute to understanding how algorithmic choices interact with modern computing architectures for applications in computer architecture, data science, and real-time embedded systems.
format Preprint
id arxiv_https___arxiv_org_abs_2606_01174
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MAC Performance and Algorithmic Optimization in Matrix Multiplication Workloads
Negasa, D
Mohamed, N. O
Ghani, Arfan
Signal Processing
Matrix multiplication is a fundamental computational kernel underlying a wide range of real-world applications, including machine learning, scientific computing, signal processing, and computer graphics. Its performance directly impacts the efficiency, scalability, and energy consumption of modern computing systems. This paper presents a comparative analysis of several matrix multiplication algorithms implemented in software and examined in the context of their hardware execution characteristics. Naive, NumPy, Strassen, and Winograd algorithms are evaluated based on execution time, user time, and CPU time across increasing matrix sizes. The performance metrics reveal computational bottlenecks and highlight the benefits of algorithmic optimizations. Furthermore, the study investigates the mathematical operations underlying each algorithm and analyzes how matrix dimensions influence MAC (Multiply-Accumulate) behavior and overall computational efficiency in the hardware domain. The results provide a performance benchmark and contribute to understanding how algorithmic choices interact with modern computing architectures for applications in computer architecture, data science, and real-time embedded systems.
title MAC Performance and Algorithmic Optimization in Matrix Multiplication Workloads
topic Signal Processing
url https://arxiv.org/abs/2606.01174