High-Rate Quantized Matrix Multiplication I
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916010031841280 |
|---|---|
| author | Ordentlich, Or Polyanskiy, Yury |
| author_facet | Ordentlich, Or Polyanskiy, Yury |
| contents | This paper investigates the problem of quantized matrix multiplication (MatMul), which has become crucial for the efficient deployment of large language models (LLMs). We consider a Generic MatMul setting, where both matrices must be quantized (weight+activation quantization) without specific apriori (calibration) statistical information about the factors. We review the fundamental information-theoretic tradeoff between quantization rate and distortion (high-rate theory), and contrast those with the performance of popular quantization schemes (absmax INT and floating-point (FP)), for which we also derive accurate heuristic approximations. Part II of this paper studies the weight-only quantization setup where second-order statistics of the activation matrices are available at the encoder. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_17187 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | High-Rate Quantized Matrix Multiplication I Ordentlich, Or Polyanskiy, Yury Information Theory Artificial Intelligence This paper investigates the problem of quantized matrix multiplication (MatMul), which has become crucial for the efficient deployment of large language models (LLMs). We consider a Generic MatMul setting, where both matrices must be quantized (weight+activation quantization) without specific apriori (calibration) statistical information about the factors. We review the fundamental information-theoretic tradeoff between quantization rate and distortion (high-rate theory), and contrast those with the performance of popular quantization schemes (absmax INT and floating-point (FP)), for which we also derive accurate heuristic approximations. Part II of this paper studies the weight-only quantization setup where second-order statistics of the activation matrices are available at the encoder. |
| title | High-Rate Quantized Matrix Multiplication I |
| topic | Information Theory Artificial Intelligence |
| url | https://arxiv.org/abs/2601.17187 |