High-Rate Quantized Matrix Multiplication I

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ordentlich, Or, Polyanskiy, Yury
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916010031841280
author Ordentlich, Or
Polyanskiy, Yury
author_facet Ordentlich, Or
Polyanskiy, Yury
contents This paper investigates the problem of quantized matrix multiplication (MatMul), which has become crucial for the efficient deployment of large language models (LLMs). We consider a Generic MatMul setting, where both matrices must be quantized (weight+activation quantization) without specific apriori (calibration) statistical information about the factors. We review the fundamental information-theoretic tradeoff between quantization rate and distortion (high-rate theory), and contrast those with the performance of popular quantization schemes (absmax INT and floating-point (FP)), for which we also derive accurate heuristic approximations. Part II of this paper studies the weight-only quantization setup where second-order statistics of the activation matrices are available at the encoder.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17187
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle High-Rate Quantized Matrix Multiplication I
Ordentlich, Or
Polyanskiy, Yury
Information Theory
Artificial Intelligence
This paper investigates the problem of quantized matrix multiplication (MatMul), which has become crucial for the efficient deployment of large language models (LLMs). We consider a Generic MatMul setting, where both matrices must be quantized (weight+activation quantization) without specific apriori (calibration) statistical information about the factors. We review the fundamental information-theoretic tradeoff between quantization rate and distortion (high-rate theory), and contrast those with the performance of popular quantization schemes (absmax INT and floating-point (FP)), for which we also derive accurate heuristic approximations. Part II of this paper studies the weight-only quantization setup where second-order statistics of the activation matrices are available at the encoder.
title High-Rate Quantized Matrix Multiplication I
topic Information Theory
Artificial Intelligence
url https://arxiv.org/abs/2601.17187