Arithmetic-Intensity-Aware Quantization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singh, Taig, Rajan, Shreshth, Jain, Nikhil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915680279855104
author Singh, Taig
Rajan, Shreshth
Jain, Nikhil
author_facet Singh, Taig
Rajan, Shreshth
Jain, Nikhil
contents As modern neural networks become increasingly memory-bound, inference throughput is limited by DRAM bandwidth rather than compute. We present Arithmetic-Intensity-Aware Quantization (AIQ), a mixed precision quantization framework that chooses per-layer bit-widths to maximize arithmetic intensity (AI) while minimizing accuracy loss. AIQ is a post-training quantization method that uses search algorithms over per-layer quantization schemes to minimize a weighted loss over AI and accuracy. On ResNet-20/CIFAR-10, AIQ increases AI by ~50% over an FP32 baseline while keeping test accuracy within ~1 percentage point, and outperforming global uniform quantization schemes. On a memory-bound MobileNetV2 architecture, AIQ configurations give a 1.66x higher throughput than the FP32 baseline while keeping test accuracy within 1 percentage point. We also find that AIQ naturally quantizes larger layers more aggressively.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14090
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Arithmetic-Intensity-Aware Quantization
Singh, Taig
Rajan, Shreshth
Jain, Nikhil
Machine Learning
Artificial Intelligence
As modern neural networks become increasingly memory-bound, inference throughput is limited by DRAM bandwidth rather than compute. We present Arithmetic-Intensity-Aware Quantization (AIQ), a mixed precision quantization framework that chooses per-layer bit-widths to maximize arithmetic intensity (AI) while minimizing accuracy loss. AIQ is a post-training quantization method that uses search algorithms over per-layer quantization schemes to minimize a weighted loss over AI and accuracy. On ResNet-20/CIFAR-10, AIQ increases AI by ~50% over an FP32 baseline while keeping test accuracy within ~1 percentage point, and outperforming global uniform quantization schemes. On a memory-bound MobileNetV2 architecture, AIQ configurations give a 1.66x higher throughput than the FP32 baseline while keeping test accuracy within 1 percentage point. We also find that AIQ naturally quantizes larger layers more aggressively.
title Arithmetic-Intensity-Aware Quantization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2512.14090