4-bit Batched GEMM on Commodity Hardware: Roofline Analysis, big.LITTLE Scheduling, and Batched Arithmetic Intensity Scaling to 243 GMACs/s on a Commercial Mobile SoC.
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Recurso digital |
| Veröffentlicht: |
Zenodo
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866901467517943808 |
|---|---|
| author | Pirolo, Andrés Sebastián |
| author_facet | Pirolo, Andrés Sebastián |
| contents | <p>We present a systematic empirical characterization of 4-bit integer matrix multiplication (INT4 GEMM) on a commercial heterogeneous mobile processor (Qualcomm Snapdragon 8 Gen 2, Samsung Galaxy Z Fold 5) using hand-written ARM NEON SIMD kernels in C++. Two INT4 values are packed per byte; 32 nibbles are processed per NEON cycle using 128-bit vector registers. All results are validated with exact integer equality against an independent scalar reference. </p> <p>To the best of the author's knowledge, this is the first systematic empirical Roofline sweep of INT4 hardware across batch sizes 1 through 32 on a commercial SoC device. For an 8192\times8192 INT4 layer, 8-core throughput at batch=1 saturates at 141.30 GMACs/s, consistent with the LPDDR5X bandwidth ceiling. Batched inference yields a peak of 243.15 GMACs/s at batch=8. Beyond batch=8, throughput degrades as the activation working set exceeds the 64 KB L1 data cache. This non-monotonic curve identifies batch=8 as the hardware-specific throughput optimum for this SoC generation, a result not previously documented in the literature. </p> <p> </p> <p><strong>This work is accompanied by a ZIP archive provided here to make the work fully self-contained. Particular attention should be given to the scope of the included license.</strong></p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19015910 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | 4-bit Batched GEMM on Commodity Hardware: Roofline Analysis, big.LITTLE Scheduling, and Batched Arithmetic Intensity Scaling to 243 GMACs/s on a Commercial Mobile SoC. Pirolo, Andrés Sebastián GEMM Edge AI mobile inference INT4 quantization Green AI <p>We present a systematic empirical characterization of 4-bit integer matrix multiplication (INT4 GEMM) on a commercial heterogeneous mobile processor (Qualcomm Snapdragon 8 Gen 2, Samsung Galaxy Z Fold 5) using hand-written ARM NEON SIMD kernels in C++. Two INT4 values are packed per byte; 32 nibbles are processed per NEON cycle using 128-bit vector registers. All results are validated with exact integer equality against an independent scalar reference. </p> <p>To the best of the author's knowledge, this is the first systematic empirical Roofline sweep of INT4 hardware across batch sizes 1 through 32 on a commercial SoC device. For an 8192\times8192 INT4 layer, 8-core throughput at batch=1 saturates at 141.30 GMACs/s, consistent with the LPDDR5X bandwidth ceiling. Batched inference yields a peak of 243.15 GMACs/s at batch=8. Beyond batch=8, throughput degrades as the activation working set exceeds the 64 KB L1 data cache. This non-monotonic curve identifies batch=8 as the hardware-specific throughput optimum for this SoC generation, a result not previously documented in the literature. </p> <p> </p> <p><strong>This work is accompanied by a ZIP archive provided here to make the work fully self-contained. Particular attention should be given to the scope of the included license.</strong></p> |
| title | 4-bit Batched GEMM on Commodity Hardware: Roofline Analysis, big.LITTLE Scheduling, and Batched Arithmetic Intensity Scaling to 243 GMACs/s on a Commercial Mobile SoC. |
| topic | GEMM Edge AI mobile inference INT4 quantization Green AI |
| url | https://doi.org/10.5281/zenodo.19015910 |