M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912853962784768 |
|---|---|
| author | Hu, Weiming Zhang, Zihan Zhang, Haoyan Zhang, Chen Guo, Cong Feng, Yu Hu, Tianchi Li, Guanglin Hu, Guipeng Wang, Junsong Leng, Jingwen |
| author_facet | Hu, Weiming Zhang, Zihan Zhang, Haoyan Zhang, Chen Guo, Cong Feng, Yu Hu, Tianchi Li, Guanglin Hu, Guipeng Wang, Junsong Leng, Jingwen |
| contents | Existing low-bit Microscaling (MX) formats, such as MXFP4, often suffer from substantial accuracy degradation due to the use of a shared scaling factor with the Power-of-Two format. In this work, we explore strategies that introduce minimal metadata to recover accuracy lost during quantization while maintaining high bit efficiency across a wide range of large language models. We propose a complete algorithm-hardware co-design based on flexible metadata, featuring an online quantization with simple encoding. To support the proposed method efficiently, we implement a lightweight hardware unit and integrate it into the accelerator. Evaluation results demonstrate that our method substantially narrows the accuracy gap, achieving on average a 70.63% reduction in accuracy loss compared to MXFP4 and a 37.30% reduction relative to the latest NVFP4 on LLM benchmarks. Furthermore, our design delivers up to 1.91$\times$ speedup and 1.75$\times$ energy savings over state-of-the-art accelerators. Our code is available at https://github.com/SJTU-ReArch-Group/M2XFP_ASPLOS26. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_19213 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization Hu, Weiming Zhang, Zihan Zhang, Haoyan Zhang, Chen Guo, Cong Feng, Yu Hu, Tianchi Li, Guanglin Hu, Guipeng Wang, Junsong Leng, Jingwen Hardware Architecture Existing low-bit Microscaling (MX) formats, such as MXFP4, often suffer from substantial accuracy degradation due to the use of a shared scaling factor with the Power-of-Two format. In this work, we explore strategies that introduce minimal metadata to recover accuracy lost during quantization while maintaining high bit efficiency across a wide range of large language models. We propose a complete algorithm-hardware co-design based on flexible metadata, featuring an online quantization with simple encoding. To support the proposed method efficiently, we implement a lightweight hardware unit and integrate it into the accelerator. Evaluation results demonstrate that our method substantially narrows the accuracy gap, achieving on average a 70.63% reduction in accuracy loss compared to MXFP4 and a 37.30% reduction relative to the latest NVFP4 on LLM benchmarks. Furthermore, our design delivers up to 1.91$\times$ speedup and 1.75$\times$ energy savings over state-of-the-art accelerators. Our code is available at https://github.com/SJTU-ReArch-Group/M2XFP_ASPLOS26. |
| title | M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization |
| topic | Hardware Architecture |
| url | https://arxiv.org/abs/2601.19213 |