Normalized Architectures are Natively 4-Bit
Fuente:
arXiv
Saved in:
| Main Authors: | Fishman, Maxim, Chmiel, Brian, Banner, Ron, Soudry, Daniel, Ginsburg, Boris |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025)
by: Chmiel, Brian, et al.
Published: (2025)
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024)
by: Fishman, Maxim, et al.
Published: (2024)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
by: Chmiel, Brian, et al.
Published: (2022)
by: Chmiel, Brian, et al.
Published: (2022)
Workspace Optimization: How to Train Your Agent
by: Sarafian, Elad, et al.
Published: (2026)
by: Sarafian, Elad, et al.
Published: (2026)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
by: Shkolnik, Moran, et al.
Published: (2024)
by: Shkolnik, Moran, et al.
Published: (2024)
Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
by: Chmiel, Brian, et al.
Published: (2021)
by: Chmiel, Brian, et al.
Published: (2021)
Retrieval from Within: An Intrinsic Capability of Attention-Based Models
by: Hoffer, Elad, et al.
Published: (2026)
by: Hoffer, Elad, et al.
Published: (2026)
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
by: Blumenfeld, Yaniv, et al.
Published: (2024)
by: Blumenfeld, Yaniv, et al.
Published: (2024)
nGPT: Normalized Transformer with Representation Learning on the Hypersphere
by: Loshchilov, Ilya, et al.
Published: (2024)
by: Loshchilov, Ilya, et al.
Published: (2024)
Learning Rate Transfer in Normalized Transformers
by: Shigida, Boris, et al.
Published: (2026)
by: Shigida, Boris, et al.
Published: (2026)
Star Attention: Efficient LLM Inference over Long Sequences
by: Acharya, Shantanu, et al.
Published: (2024)
by: Acharya, Shantanu, et al.
Published: (2024)
On Exact Bit-level Reversible Transformers Without Changing Architectures
by: Zhang, Guoqiang, et al.
Published: (2024)
by: Zhang, Guoqiang, et al.
Published: (2024)
Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step Sizes
by: Qiao, Dan, et al.
Published: (2024)
by: Qiao, Dan, et al.
Published: (2024)
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
by: Son, Donghyun, et al.
Published: (2025)
by: Son, Donghyun, et al.
Published: (2025)
Human-aligned Chess with a Bit of Search
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
by: Zhang, Peiyuan, et al.
Published: (2026)
by: Zhang, Peiyuan, et al.
Published: (2026)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
by: Lee, Deokjae, et al.
Published: (2025)
by: Lee, Deokjae, et al.
Published: (2025)
Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
by: Ficek, Aleksander, et al.
Published: (2025)
by: Ficek, Aleksander, et al.
Published: (2025)
Pretraining large language models with MXFP4 on Native FP4 Hardware
by: Cim, Musa, et al.
Published: (2026)
by: Cim, Musa, et al.
Published: (2026)
BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization
by: Zhao, Jiayu, et al.
Published: (2026)
by: Zhao, Jiayu, et al.
Published: (2026)
Tiled Bit Networks: Sub-Bit Neural Network Compression Through Reuse of Learnable Binary Vectors
by: Gorbett, Matt, et al.
Published: (2024)
by: Gorbett, Matt, et al.
Published: (2024)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
by: Becking, Daniel, et al.
Published: (2021)
by: Becking, Daniel, et al.
Published: (2021)
XGRAG: A Graph-Native Framework for Explaining KG-based Retrieval-Augmented Generation
by: Li, Zhuoling, et al.
Published: (2026)
by: Li, Zhuoling, et al.
Published: (2026)
CrashFormer: A Multimodal Architecture to Predict the Risk of Crash
by: Monsefi, Amin Karimi, et al.
Published: (2024)
by: Monsefi, Amin Karimi, et al.
Published: (2024)
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
by: Kim, Jeonghoon, et al.
Published: (2025)
by: Kim, Jeonghoon, et al.
Published: (2025)
Scaling CrossQ with Weight Normalization
by: Palenicek, Daniel, et al.
Published: (2025)
by: Palenicek, Daniel, et al.
Published: (2025)
LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment
by: Lee, Banseok, et al.
Published: (2026)
by: Lee, Banseok, et al.
Published: (2026)
Label-Looping: Highly Efficient Decoding for Transducers
by: Bataev, Vladimir, et al.
Published: (2024)
by: Bataev, Vladimir, et al.
Published: (2024)
Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
by: Zhao, Maosen, et al.
Published: (2025)
by: Zhao, Maosen, et al.
Published: (2025)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
by: Lee, Banseok, et al.
Published: (2025)
by: Lee, Banseok, et al.
Published: (2025)
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
by: Feng, Jingyuan, et al.
Published: (2026)
by: Feng, Jingyuan, et al.
Published: (2026)
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
by: Gernigon, Cédric, et al.
Published: (2024)
by: Gernigon, Cédric, et al.
Published: (2024)
Unsupervised Representation Learning - an Invariant Risk Minimization Perspective
by: Norman, Yotam, et al.
Published: (2025)
by: Norman, Yotam, et al.
Published: (2025)
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
by: Palenicek, Daniel, et al.
Published: (2025)
by: Palenicek, Daniel, et al.
Published: (2025)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
Bit-Identical Medical Deep Learning via Structured Orthogonal Initialization
by: Shkolnikov, Yakov Pyotr
Published: (2026)
by: Shkolnikov, Yakov Pyotr
Published: (2026)
Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
by: Zhang, Xi, et al.
Published: (2025)
by: Zhang, Xi, et al.
Published: (2025)
Why Do Some Inputs Break Low-Bit LLM Quantization?
by: Chang, Ting-Yun, et al.
Published: (2025)
by: Chang, Ting-Yun, et al.
Published: (2025)
Pushing the Limits of Low-Bit Optimizers: A Focus on EMA Dynamics
by: Xu, Cong, et al.
Published: (2025)
by: Xu, Cong, et al.
Published: (2025)
HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs
by: Chen, Ningning, et al.
Published: (2025)
by: Chen, Ningning, et al.
Published: (2025)
Similar Items
-
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025) -
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024) -
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
by: Chmiel, Brian, et al.
Published: (2022) -
Workspace Optimization: How to Train Your Agent
by: Sarafian, Elad, et al.
Published: (2026) -
EXAQ: Exponent Aware Quantization For LLMs Acceleration
by: Shkolnik, Moran, et al.
Published: (2024)