MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Jungi, Park, Junyong, Cha, Soohyun, Cho, Jaehoon, Sim, Jaewoong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915557837635584
author Lee, Jungi
Park, Junyong
Cha, Soohyun
Cho, Jaehoon
Sim, Jaewoong
author_facet Lee, Jungi
Park, Junyong
Cha, Soohyun
Cho, Jaehoon
Sim, Jaewoong
contents Reduced-precision data formats are crucial for cost-effective serving of large language models (LLMs). While numerous reduced-precision formats have been introduced thus far, they often require intrusive modifications to the software frameworks or are rather unconventional for widespread adoption across hardware vendors. In this paper, we instead focus on recent industry-driven variants of block floating-point (BFP) formats and conduct a comprehensive analysis to push their limits for efficient LLM serving. Our analysis shows that existing ultra low-bit BFP variants struggle to provide reasonable language model performance due to outlier values in blocks. To address the outliers with BFPs, we propose MX+, a cost-effective and non-intrusive extension designed for seamless integration into the microscaling (MX) formats. MX+ builds on the key insight that the outlier does not need to use its exponent field in the element data type, which allows us to repurpose the exponent field as an extended mantissa to increase the precision of the outlier element. Our evaluation shows that MX+ achieves significantly higher model performance compared to the 4-bit MX format (MXFP4) with negligible storage overhead and slowdown, thus offering a compelling alternative to MXFP4 or MXFP6 for efficient LLM inference.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14557
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
Lee, Jungi
Park, Junyong
Cha, Soohyun
Cho, Jaehoon
Sim, Jaewoong
Machine Learning
Hardware Architecture
Reduced-precision data formats are crucial for cost-effective serving of large language models (LLMs). While numerous reduced-precision formats have been introduced thus far, they often require intrusive modifications to the software frameworks or are rather unconventional for widespread adoption across hardware vendors. In this paper, we instead focus on recent industry-driven variants of block floating-point (BFP) formats and conduct a comprehensive analysis to push their limits for efficient LLM serving. Our analysis shows that existing ultra low-bit BFP variants struggle to provide reasonable language model performance due to outlier values in blocks. To address the outliers with BFPs, we propose MX+, a cost-effective and non-intrusive extension designed for seamless integration into the microscaling (MX) formats. MX+ builds on the key insight that the outlier does not need to use its exponent field in the element data type, which allows us to repurpose the exponent field as an extended mantissa to increase the precision of the outlier element. Our evaluation shows that MX+ achieves significantly higher model performance compared to the 4-bit MX format (MXFP4) with negligible storage overhead and slowdown, thus offering a compelling alternative to MXFP4 or MXFP6 for efficient LLM inference.
title MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
topic Machine Learning
Hardware Architecture
url https://arxiv.org/abs/2510.14557