HiFloat4 Format for Language Model Inference

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Luo, Yuanyong, Huang, Jing, Cheng, Yu, Yu, Ziwei, Tang, Kaihua, Ma, Xinda, Wang, Xin, Tong, Anping, Hu, Guipeng, Xu, Yun, Taghian, Mehran, Wu, Peng, Li, Guanglin, Peng, Yunke, Hu, Tianchi, Chen, Minqi, Mi, Michael Bi, Liu, Hu, Zhou, Xiping, Wang, Junsong, Lin, Qiang, Liao, Heng
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911444314882048
author Luo, Yuanyong
Huang, Jing
Cheng, Yu
Yu, Ziwei
Tang, Kaihua
Ma, Xinda
Wang, Xin
Tong, Anping
Hu, Guipeng
Xu, Yun
Taghian, Mehran
Wu, Peng
Li, Guanglin
Peng, Yunke
Hu, Tianchi
Chen, Minqi
Mi, Michael Bi
Liu, Hu
Zhou, Xiping
Wang, Junsong
Lin, Qiang
Liao, Heng
author_facet Luo, Yuanyong
Huang, Jing
Cheng, Yu
Yu, Ziwei
Tang, Kaihua
Ma, Xinda
Wang, Xin
Tong, Anping
Hu, Guipeng
Xu, Yun
Taghian, Mehran
Wu, Peng
Li, Guanglin
Peng, Yunke
Hu, Tianchi
Chen, Minqi
Mi, Michael Bi
Liu, Hu
Zhou, Xiping
Wang, Junsong
Lin, Qiang
Liao, Heng
contents This paper introduces HiFloat4 (HiF4), a block floating-point data format tailored for deep learning. Each HiF4 unit packs 64 4-bit elements with 32 bits of shared scaling metadata, averaging 4.5 bits per value. The metadata specifies a three-level scaling hierarchy, capturing inter- and intra-group dynamic range while improving the utilization of the representational space. In addition, the large 64-element group size enables matrix multiplications to be executed in a highly fixed-point manner, significantly reducing hardware area and power consumption. To evaluate the proposed format, we conducted inference experiments on several language models, including LLaMA, Qwen, Mistral, DeepSeek-V3.1 and LongCat. Results show that HiF4 achieves higher average accuracy than the state-of-the-art NVFP4 format across multiple models and diverse downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_11287
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HiFloat4 Format for Language Model Inference
Luo, Yuanyong
Huang, Jing
Cheng, Yu
Yu, Ziwei
Tang, Kaihua
Ma, Xinda
Wang, Xin
Tong, Anping
Hu, Guipeng
Xu, Yun
Taghian, Mehran
Wu, Peng
Li, Guanglin
Peng, Yunke
Hu, Tianchi
Chen, Minqi
Mi, Michael Bi
Liu, Hu
Zhou, Xiping
Wang, Junsong
Lin, Qiang
Liao, Heng
Machine Learning
Artificial Intelligence
Hardware Architecture
This paper introduces HiFloat4 (HiF4), a block floating-point data format tailored for deep learning. Each HiF4 unit packs 64 4-bit elements with 32 bits of shared scaling metadata, averaging 4.5 bits per value. The metadata specifies a three-level scaling hierarchy, capturing inter- and intra-group dynamic range while improving the utilization of the representational space. In addition, the large 64-element group size enables matrix multiplications to be executed in a highly fixed-point manner, significantly reducing hardware area and power consumption. To evaluate the proposed format, we conducted inference experiments on several language models, including LLaMA, Qwen, Mistral, DeepSeek-V3.1 and LongCat. Results show that HiF4 achieves higher average accuracy than the state-of-the-art NVFP4 format across multiple models and diverse downstream tasks.
title HiFloat4 Format for Language Model Inference
topic Machine Learning
Artificial Intelligence
Hardware Architecture
url https://arxiv.org/abs/2602.11287