Is Finer Better? The Limits of Microscaling Formats in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fasoli, Andrea, Kar, Monodeep, Liu, Chi-Chun, Venkataramani, Swagath, Srinivasan, Viji, Chang, Leland, Wang, Naigang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917224919334912
author Fasoli, Andrea
Kar, Monodeep
Liu, Chi-Chun
Venkataramani, Swagath
Srinivasan, Viji
Chang, Leland
Wang, Naigang
author_facet Fasoli, Andrea
Kar, Monodeep
Liu, Chi-Chun
Venkataramani, Swagath
Srinivasan, Viji
Chang, Leland
Wang, Naigang
contents Microscaling data formats leverage per-block tensor quantization to enable aggressive model compression with limited loss in accuracy. Unlocking their potential for efficient training and inference necessitates hardware-friendly implementations that handle matrix multiplications in a native format and adopt efficient error-mitigation strategies. Herein, we report the emergence of a surprising behavior associated with microscaling quantization, whereas the output of a quantized model degrades as block size is decreased below a given threshold. This behavior clashes with the expectation that a smaller block size should allow for a better representation of the tensor elements. We investigate this phenomenon both experimentally and theoretically, decoupling the sources of quantization error behind it. Experimentally, we analyze the distributions of several Large Language Models and identify the conditions driving the anomalous behavior. Theoretically, we lay down a framework showing remarkable agreement with experimental data from pretrained model distributions and ideal ones. Overall, we show that the anomaly is driven by the interplay between narrow tensor distributions and the limited dynamic range of the quantized scales. Based on these insights, we propose the use of FP8 unsigned E5M3 (UE5M3) as a novel hardware-friendly format for the scales in FP4 microscaling data types. We demonstrate that UE5M3 achieves comparable performance to the conventional FP8 unsigned E4M3 scales while obviating the need of global scaling operations on weights and activations.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19026
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Is Finer Better? The Limits of Microscaling Formats in Large Language Models
Fasoli, Andrea
Kar, Monodeep
Liu, Chi-Chun
Venkataramani, Swagath
Srinivasan, Viji
Chang, Leland
Wang, Naigang
Machine Learning
Hardware Architecture
Computation and Language
I.2.6
Microscaling data formats leverage per-block tensor quantization to enable aggressive model compression with limited loss in accuracy. Unlocking their potential for efficient training and inference necessitates hardware-friendly implementations that handle matrix multiplications in a native format and adopt efficient error-mitigation strategies. Herein, we report the emergence of a surprising behavior associated with microscaling quantization, whereas the output of a quantized model degrades as block size is decreased below a given threshold. This behavior clashes with the expectation that a smaller block size should allow for a better representation of the tensor elements. We investigate this phenomenon both experimentally and theoretically, decoupling the sources of quantization error behind it. Experimentally, we analyze the distributions of several Large Language Models and identify the conditions driving the anomalous behavior. Theoretically, we lay down a framework showing remarkable agreement with experimental data from pretrained model distributions and ideal ones. Overall, we show that the anomaly is driven by the interplay between narrow tensor distributions and the limited dynamic range of the quantized scales. Based on these insights, we propose the use of FP8 unsigned E5M3 (UE5M3) as a novel hardware-friendly format for the scales in FP4 microscaling data types. We demonstrate that UE5M3 achieves comparable performance to the conventional FP8 unsigned E4M3 scales while obviating the need of global scaling operations on weights and activations.
title Is Finer Better? The Limits of Microscaling Formats in Large Language Models
topic Machine Learning
Hardware Architecture
Computation and Language
I.2.6
url https://arxiv.org/abs/2601.19026