Normalized Architectures are Natively 4-Bit

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fishman, Maxim, Chmiel, Brian, Banner, Ron, Soudry, Daniel, Ginsburg, Boris
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914539422875648
author Fishman, Maxim
Chmiel, Brian
Banner, Ron
Soudry, Daniel
Ginsburg, Boris
author_facet Fishman, Maxim
Chmiel, Brian
Banner, Ron
Soudry, Daniel
Ginsburg, Boris
contents Training large language models at 4-bit precision is critical for efficiency. We show that nGPT, an architecture that constrains weights and hidden representations to the unit hypersphere, is inherently more robust to low-precision arithmetic. This removes the need for interventions-such as applying random Hadamard transforms and performing per-tensor scaling calculations-to preserve model quality, and it enables stable end-to-end NVFP4 training. We validate this approach on both a 1.2B dense model and hybrid (Mamba-Transformer) MoE models of up to 3B/30B parameters. We trace this robustness to the dot product: while quantization noise remains largely uncorrelated in both standard and normalized architectures, the signal behaves differently. In nGPT, the hypersphere constraint enhances weak positive correlations among the element-wise products, leading to a constructive accumulation of the signal across the hidden dimension while the noise continues to average out. This yields a higher effective signal-to-noise ratio and a flatter loss landscape, with the effect strengthening as the hidden dimension grows, suggesting increasing advantages at scale. A reference implementation is available at https://github.com/anonymous452026/ngpt-nvfp4
format Preprint
id arxiv_https___arxiv_org_abs_2605_06067
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Normalized Architectures are Natively 4-Bit
Fishman, Maxim
Chmiel, Brian
Banner, Ron
Soudry, Daniel
Ginsburg, Boris
Machine Learning
Artificial Intelligence
Training large language models at 4-bit precision is critical for efficiency. We show that nGPT, an architecture that constrains weights and hidden representations to the unit hypersphere, is inherently more robust to low-precision arithmetic. This removes the need for interventions-such as applying random Hadamard transforms and performing per-tensor scaling calculations-to preserve model quality, and it enables stable end-to-end NVFP4 training. We validate this approach on both a 1.2B dense model and hybrid (Mamba-Transformer) MoE models of up to 3B/30B parameters. We trace this robustness to the dot product: while quantization noise remains largely uncorrelated in both standard and normalized architectures, the signal behaves differently. In nGPT, the hypersphere constraint enhances weak positive correlations among the element-wise products, leading to a constructive accumulation of the signal across the hidden dimension while the noise continues to average out. This yields a higher effective signal-to-noise ratio and a flatter loss landscape, with the effect strengthening as the hidden dimension grows, suggesting increasing advantages at scale. A reference implementation is available at https://github.com/anonymous452026/ngpt-nvfp4
title Normalized Architectures are Natively 4-Bit
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.06067