4-bit Batched GEMM on Commodity Hardware: Roofline Analysis, big.LITTLE Scheduling, and Batched Arithmetic Intensity Scaling to 243 GMACs/s on a Commercial Mobile SoC.

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Pirolo, Andrés Sebastián
Format: Recurso digital
Veröffentlicht: Zenodo 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901467517943808
author Pirolo, Andrés Sebastián
author_facet Pirolo, Andrés Sebastián
contents <p>We present a systematic empirical characterization of 4-bit integer matrix multiplication (INT4 GEMM) on a commercial heterogeneous mobile processor (Qualcomm Snapdragon 8 Gen 2, Samsung Galaxy Z Fold 5) using hand-written ARM NEON SIMD kernels in C++. Two INT4 values are packed per byte; 32 nibbles are processed per NEON cycle using 128-bit vector registers. All results are validated with exact integer equality against an independent scalar reference.  </p> <p>To the best of the author's knowledge, this is the first systematic empirical Roofline sweep of INT4 hardware across batch sizes 1 through 32 on a commercial SoC device. For an 8192\times8192 INT4 layer, 8-core throughput at batch=1 saturates at 141.30 GMACs/s, consistent with the LPDDR5X bandwidth ceiling. Batched inference yields a peak of 243.15 GMACs/s at batch=8. Beyond batch=8, throughput degrades as the activation working set exceeds the 64 KB L1 data cache. This non-monotonic curve identifies batch=8 as the hardware-specific throughput optimum for this SoC generation, a result not previously documented in the literature. </p> <p> </p> <p><strong>This work is accompanied by a ZIP archive provided here to make the work fully self-contained. Particular attention should be given to the scope of the included license.</strong></p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19015910
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle 4-bit Batched GEMM on Commodity Hardware: Roofline Analysis, big.LITTLE Scheduling, and Batched Arithmetic Intensity Scaling to 243 GMACs/s on a Commercial Mobile SoC.
Pirolo, Andrés Sebastián
GEMM
Edge AI
mobile inference
INT4 quantization
Green AI
<p>We present a systematic empirical characterization of 4-bit integer matrix multiplication (INT4 GEMM) on a commercial heterogeneous mobile processor (Qualcomm Snapdragon 8 Gen 2, Samsung Galaxy Z Fold 5) using hand-written ARM NEON SIMD kernels in C++. Two INT4 values are packed per byte; 32 nibbles are processed per NEON cycle using 128-bit vector registers. All results are validated with exact integer equality against an independent scalar reference.  </p> <p>To the best of the author's knowledge, this is the first systematic empirical Roofline sweep of INT4 hardware across batch sizes 1 through 32 on a commercial SoC device. For an 8192\times8192 INT4 layer, 8-core throughput at batch=1 saturates at 141.30 GMACs/s, consistent with the LPDDR5X bandwidth ceiling. Batched inference yields a peak of 243.15 GMACs/s at batch=8. Beyond batch=8, throughput degrades as the activation working set exceeds the 64 KB L1 data cache. This non-monotonic curve identifies batch=8 as the hardware-specific throughput optimum for this SoC generation, a result not previously documented in the literature. </p> <p> </p> <p><strong>This work is accompanied by a ZIP archive provided here to make the work fully self-contained. Particular attention should be given to the scope of the included license.</strong></p>
title 4-bit Batched GEMM on Commodity Hardware: Roofline Analysis, big.LITTLE Scheduling, and Batched Arithmetic Intensity Scaling to 243 GMACs/s on a Commercial Mobile SoC.
topic GEMM
Edge AI
mobile inference
INT4 quantization
Green AI
url https://doi.org/10.5281/zenodo.19015910