HiFloat4 Format for Language Model Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Yuanyong, Huang, Jing, Cheng, Yu, Yu, Ziwei, Tang, Kaihua, Ma, Xinda, Wang, Xin, Tong, Anping, Hu, Guipeng, Xu, Yun, Taghian, Mehran, Wu, Peng, Li, Guanglin, Peng, Yunke, Hu, Tianchi, Chen, Minqi, Mi, Michael Bi, Liu, Hu, Zhou, Xiping, Wang, Junsong, Lin, Qiang, Liao, Heng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ascend HiFloat8 Format for Deep Learning
by: Luo, Yuanyong, et al.
Published: (2024)
by: Luo, Yuanyong, et al.
Published: (2024)
HiFloat4 Format for Language Model Pre-training on Ascend NPUs
by: Taghian, Mehran, et al.
Published: (2026)
by: Taghian, Mehran, et al.
Published: (2026)
M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
by: Hu, Weiming, et al.
Published: (2026)
by: Hu, Weiming, et al.
Published: (2026)
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
by: Zhao, Pengxiang, et al.
Published: (2026)
by: Zhao, Pengxiang, et al.
Published: (2026)
From Quarter to All: Accelerating Speculative LLM Decoding via Floating-Point Exponent Remapping and Parameter Sharing
by: Zhao, Yushu, et al.
Published: (2025)
by: Zhao, Yushu, et al.
Published: (2025)
A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators
by: Li, Cong, et al.
Published: (2026)
by: Li, Cong, et al.
Published: (2026)
Apple vs. Oranges: Evaluating the Apple Silicon M-Series SoCs for HPC Performance and Efficiency
by: Hübner, Paul, et al.
Published: (2025)
by: Hübner, Paul, et al.
Published: (2025)
Fast Cross-Operator Optimization of Attention Dataflow
by: Chang, Haodong, et al.
Published: (2026)
by: Chang, Haodong, et al.
Published: (2026)
TimeFloats: Train-in-Memory with Time-Domain Floating-Point Scalar Products
by: Hashem, Maeesha Binte, et al.
Published: (2024)
by: Hashem, Maeesha Binte, et al.
Published: (2024)
BBAL: A Bidirectional Block Floating Point-Based Quantisation Accelerator for Large Language Models
by: Han, Xiaomeng, et al.
Published: (2025)
by: Han, Xiaomeng, et al.
Published: (2025)
HiHGNN: Accelerating HGNNs through Parallelism and Data Reusability Exploitation
by: Xue, Runzhen, et al.
Published: (2023)
by: Xue, Runzhen, et al.
Published: (2023)
FPGA-based Emulation and Device-Side Management for CXL-based Memory Tiering Systems
by: Chen, Yiqi, et al.
Published: (2025)
by: Chen, Yiqi, et al.
Published: (2025)
Closing the Gap Between Float and Posit Hardware Efficiency
by: Jonnalagadda, Aditya Anirudh, et al.
Published: (2026)
by: Jonnalagadda, Aditya Anirudh, et al.
Published: (2026)
BoolE: Exact Symbolic Reasoning via Boolean Equality Saturation
by: Yin, Jiaqi, et al.
Published: (2025)
by: Yin, Jiaqi, et al.
Published: (2025)
Timing-driven Approximate Logic Synthesis Based on Double-chase Grey Wolf Optimizer
by: Hu, Xiangfei, et al.
Published: (2024)
by: Hu, Xiangfei, et al.
Published: (2024)
Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM
by: Yu, Zhongkai, et al.
Published: (2024)
by: Yu, Zhongkai, et al.
Published: (2024)
Scaling Laws for Floating Point Quantization Training
by: Sun, Xingwu, et al.
Published: (2025)
by: Sun, Xingwu, et al.
Published: (2025)
Fast Generation of Custom Floating-Point Spatial Filters on FPGAs
by: Campos, Nelson, et al.
Published: (2024)
by: Campos, Nelson, et al.
Published: (2024)
Online Alignment and Addition in Multi-Term Floating-Point Adders
by: Alexandridis, Kosmas, et al.
Published: (2024)
by: Alexandridis, Kosmas, et al.
Published: (2024)
Floating Point HUB Adder for RISC-V Sargantana Processor
by: Bandera, Gerardo, et al.
Published: (2024)
by: Bandera, Gerardo, et al.
Published: (2024)
LLM-Powered Code Analysis and Optimization for Gaussian Splatting Kernels
by: Hu, Yi, et al.
Published: (2025)
by: Hu, Yi, et al.
Published: (2025)
QiMeng-CPU-v2: Automated Superscalar Processor Design by Learning Data Dependencies
by: Cheng, Shuyao, et al.
Published: (2025)
by: Cheng, Shuyao, et al.
Published: (2025)
TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments
by: Guan, Yue, et al.
Published: (2026)
by: Guan, Yue, et al.
Published: (2026)
HiMA: Hierarchical Quantum Microarchitecture for Qubit-Scaling and Quantum Process-Level Parallelism
by: Zhou, Qi, et al.
Published: (2024)
by: Zhou, Qi, et al.
Published: (2024)
TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines
by: Wang, Jiayi, et al.
Published: (2026)
by: Wang, Jiayi, et al.
Published: (2026)
The AetherFloat Family: Block-Scale-Free Quad-Radix Floating-Point Architectures for AI Accelerators
by: Morisaki, Keita
Published: (2026)
by: Morisaki, Keita
Published: (2026)
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
by: Alexandridis, Kosmas, et al.
Published: (2024)
by: Alexandridis, Kosmas, et al.
Published: (2024)
Converting Binary Floating-Point Numbers to Shortest Decimal Strings: An Experimental Review
by: Gareau, Jaël Champagne, et al.
Published: (2026)
by: Gareau, Jaël Champagne, et al.
Published: (2026)
Accelerator-assisted Floating-point ASIP for Communication and Positioning in Massive MIMO Systems
by: Attari, Mohammad, et al.
Published: (2025)
by: Attari, Mohammad, et al.
Published: (2025)
Mixed Structural Choice Operator: Enhancing Technology Mapping with Heterogeneous Representations
by: Hu, Zhang, et al.
Published: (2025)
by: Hu, Zhang, et al.
Published: (2025)
Non-Overlapping Placement of Macro Cells based on Reinforcement Learning in Chip Design
by: Yu, Tao, et al.
Published: (2024)
by: Yu, Tao, et al.
Published: (2024)
A Stochastic Rounding-Enabled Low-Precision Floating-Point MAC for DNN Training
by: Ali, Sami Ben, et al.
Published: (2024)
by: Ali, Sami Ben, et al.
Published: (2024)
Search Your Block Floating Point Scales!
by: Gupta, Tanmaey, et al.
Published: (2026)
by: Gupta, Tanmaey, et al.
Published: (2026)
AGON: Automated Design Framework for Customizing Processors from ISA Documents
by: Li, Chongxiao, et al.
Published: (2024)
by: Li, Chongxiao, et al.
Published: (2024)
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
MXFormer: A Microscaling Floating-Point Charge-Trap Transistor Compute-in-Memory Transformer Accelerator
by: Karfakis, George, et al.
Published: (2026)
by: Karfakis, George, et al.
Published: (2026)
Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training
by: Nikolić, Miloš, et al.
Published: (2022)
by: Nikolić, Miloš, et al.
Published: (2022)
Theseus: Exploring Efficient Wafer-Scale Chip Design for Large Language Models
by: Zhu, Jingchen, et al.
Published: (2024)
by: Zhu, Jingchen, et al.
Published: (2024)
AI-Powered Agile Analog Circuit Design and Optimization
by: Hu, Jinhai, et al.
Published: (2025)
by: Hu, Jinhai, et al.
Published: (2025)
A Computing-in-Memory-based One-Class Hyperdimensional Computing Model for Outlier Detection
by: Wang, Ruixuan, et al.
Published: (2023)
by: Wang, Ruixuan, et al.
Published: (2023)
Similar Items
-
Ascend HiFloat8 Format for Deep Learning
by: Luo, Yuanyong, et al.
Published: (2024) -
HiFloat4 Format for Language Model Pre-training on Ascend NPUs
by: Taghian, Mehran, et al.
Published: (2026) -
M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
by: Hu, Weiming, et al.
Published: (2026) -
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
by: Zhao, Pengxiang, et al.
Published: (2026) -
From Quarter to All: Accelerating Speculative LLM Decoding via Floating-Point Exponent Remapping and Parameter Sharing
by: Zhao, Yushu, et al.
Published: (2025)