M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Weiming, Zhang, Zihan, Zhang, Haoyan, Zhang, Chen, Guo, Cong, Feng, Yu, Hu, Tianchi, Li, Guanglin, Hu, Guipeng, Wang, Junsong, Leng, Jingwen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912853962784768
author Hu, Weiming
Zhang, Zihan
Zhang, Haoyan
Zhang, Chen
Guo, Cong
Feng, Yu
Hu, Tianchi
Li, Guanglin
Hu, Guipeng
Wang, Junsong
Leng, Jingwen
author_facet Hu, Weiming
Zhang, Zihan
Zhang, Haoyan
Zhang, Chen
Guo, Cong
Feng, Yu
Hu, Tianchi
Li, Guanglin
Hu, Guipeng
Wang, Junsong
Leng, Jingwen
contents Existing low-bit Microscaling (MX) formats, such as MXFP4, often suffer from substantial accuracy degradation due to the use of a shared scaling factor with the Power-of-Two format. In this work, we explore strategies that introduce minimal metadata to recover accuracy lost during quantization while maintaining high bit efficiency across a wide range of large language models. We propose a complete algorithm-hardware co-design based on flexible metadata, featuring an online quantization with simple encoding. To support the proposed method efficiently, we implement a lightweight hardware unit and integrate it into the accelerator. Evaluation results demonstrate that our method substantially narrows the accuracy gap, achieving on average a 70.63% reduction in accuracy loss compared to MXFP4 and a 37.30% reduction relative to the latest NVFP4 on LLM benchmarks. Furthermore, our design delivers up to 1.91$\times$ speedup and 1.75$\times$ energy savings over state-of-the-art accelerators. Our code is available at https://github.com/SJTU-ReArch-Group/M2XFP_ASPLOS26.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19213
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
Hu, Weiming
Zhang, Zihan
Zhang, Haoyan
Zhang, Chen
Guo, Cong
Feng, Yu
Hu, Tianchi
Li, Guanglin
Hu, Guipeng
Wang, Junsong
Leng, Jingwen
Hardware Architecture
Existing low-bit Microscaling (MX) formats, such as MXFP4, often suffer from substantial accuracy degradation due to the use of a shared scaling factor with the Power-of-Two format. In this work, we explore strategies that introduce minimal metadata to recover accuracy lost during quantization while maintaining high bit efficiency across a wide range of large language models. We propose a complete algorithm-hardware co-design based on flexible metadata, featuring an online quantization with simple encoding. To support the proposed method efficiently, we implement a lightweight hardware unit and integrate it into the accelerator. Evaluation results demonstrate that our method substantially narrows the accuracy gap, achieving on average a 70.63% reduction in accuracy loss compared to MXFP4 and a 37.30% reduction relative to the latest NVFP4 on LLM benchmarks. Furthermore, our design delivers up to 1.91$\times$ speedup and 1.75$\times$ energy savings over state-of-the-art accelerators. Our code is available at https://github.com/SJTU-ReArch-Group/M2XFP_ASPLOS26.
title M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
topic Hardware Architecture
url https://arxiv.org/abs/2601.19213