Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
Fuente:
arXiv
Saved in:
| Main Authors: | Chmiel, Brian, Banner, Ron, Hoffer, Elad, Yaacov, Hilla Ben, Soudry, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025)
by: Chmiel, Brian, et al.
Published: (2025)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
by: Chmiel, Brian, et al.
Published: (2022)
by: Chmiel, Brian, et al.
Published: (2022)
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024)
by: Fishman, Maxim, et al.
Published: (2024)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
by: Shkolnik, Moran, et al.
Published: (2024)
by: Shkolnik, Moran, et al.
Published: (2024)
Normalized Architectures are Natively 4-Bit
by: Fishman, Maxim, et al.
Published: (2026)
by: Fishman, Maxim, et al.
Published: (2026)
Retrieval from Within: An Intrinsic Capability of Attention-Based Models
by: Hoffer, Elad, et al.
Published: (2026)
by: Hoffer, Elad, et al.
Published: (2026)
Workspace Optimization: How to Train Your Agent
by: Sarafian, Elad, et al.
Published: (2026)
by: Sarafian, Elad, et al.
Published: (2026)
The Implicit Bias of Gradient Descent on Separable Data
by: Soudry, Daniel, et al.
Published: (2017)
by: Soudry, Daniel, et al.
Published: (2017)
Distributed Training under Packet Loss
by: Weintraub, Erez, et al.
Published: (2025)
by: Weintraub, Erez, et al.
Published: (2025)
PLUMAGE: Probabilistic Low rank Unbiased Min Variance Gradient Estimator for Efficient Large Model Training
by: Haroush, Matan, et al.
Published: (2025)
by: Haroush, Matan, et al.
Published: (2025)
Optimal L2 Regularization in High-dimensional Continual Linear Regression
by: Karpel, Gilad, et al.
Published: (2026)
by: Karpel, Gilad, et al.
Published: (2026)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
by: Yan, Xianglong, et al.
Published: (2026)
by: Yan, Xianglong, et al.
Published: (2026)
Stabilizing Backpropagation in 16-bit Neural Training with Modified Adam Optimizer
by: Yun, Juyoung
Published: (2023)
by: Yun, Juyoung
Published: (2023)
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
by: Zhao, Yilong, et al.
Published: (2023)
by: Zhao, Yilong, et al.
Published: (2023)
4-bit Shampoo for Memory-Efficient Network Training
by: Wang, Sike, et al.
Published: (2024)
by: Wang, Sike, et al.
Published: (2024)
MergeQuant: Accurate 4-bit Static Quantization of Large Language Models by Channel-wise Calibration
by: Wang, Jinguang, et al.
Published: (2025)
by: Wang, Jinguang, et al.
Published: (2025)
OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization
by: Li, Zhikai, et al.
Published: (2026)
by: Li, Zhikai, et al.
Published: (2026)
When Diffusion Models Memorize: Inductive Biases in Probability Flow of Minimum-Norm Shallow Neural Nets
by: Zeno, Chen, et al.
Published: (2025)
by: Zeno, Chen, et al.
Published: (2025)
Accuracy Booster: Enabling 4-bit Fixed-point Arithmetic for DNN Training
by: Harma, Simla Burcu, et al.
Published: (2022)
by: Harma, Simla Burcu, et al.
Published: (2022)
Explore to Generalize in Zero-Shot RL
by: Zisselman, Ev, et al.
Published: (2023)
by: Zisselman, Ev, et al.
Published: (2023)
The Implicit Bias of Gradient Descent on Separable Multiclass Data
by: Ravi, Hrithik, et al.
Published: (2024)
by: Ravi, Hrithik, et al.
Published: (2024)
Foldable SuperNets: Scalable Merging of Transformers with Different Initializations and Tasks
by: Kinderman, Edan, et al.
Published: (2024)
by: Kinderman, Edan, et al.
Published: (2024)
How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers
by: Buzaglo, Gon, et al.
Published: (2024)
by: Buzaglo, Gon, et al.
Published: (2024)
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
by: Blumenfeld, Yaniv, et al.
Published: (2024)
by: Blumenfeld, Yaniv, et al.
Published: (2024)
Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
by: Li, Jinhao, et al.
Published: (2023)
by: Li, Jinhao, et al.
Published: (2023)
The Joint Effect of Task Similarity and Overparameterization on Catastrophic Forgetting -- An Analytical Model
by: Goldfarb, Daniel, et al.
Published: (2024)
by: Goldfarb, Daniel, et al.
Published: (2024)
Dynamic Rank Adjustment for Accurate and Efficient Neural Network Training
by: Shin, Hyuntak, et al.
Published: (2025)
by: Shin, Hyuntak, et al.
Published: (2025)
SILO: Solving Inverse Problems with Latent Operators
by: Raphaeli, Ron, et al.
Published: (2025)
by: Raphaeli, Ron, et al.
Published: (2025)
A Unified Characterization of Private Learnability via Graph Theory
by: Alon, Noga, et al.
Published: (2023)
by: Alon, Noga, et al.
Published: (2023)
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
by: Xia, Haojun, et al.
Published: (2025)
by: Xia, Haojun, et al.
Published: (2025)
Accurate and Scalable Matrix Mechanisms via Divide and Conquer
by: He, Guanlin, et al.
Published: (2026)
by: He, Guanlin, et al.
Published: (2026)
BitNet a4.8: 4-bit Activations for 1-bit LLMs
by: Wang, Hongyu, et al.
Published: (2024)
by: Wang, Hongyu, et al.
Published: (2024)
Tensor-Parallelism with Partially Synchronized Activations
by: Lamprecht, Itay, et al.
Published: (2025)
by: Lamprecht, Itay, et al.
Published: (2025)
How do Minimum-Norm Shallow Denoisers Look in Function Space?
by: Zeno, Chen, et al.
Published: (2023)
by: Zeno, Chen, et al.
Published: (2023)
Temperature is All You Need for Generalization in Langevin Dynamics and other Markov Processes
by: Harel, Itamar, et al.
Published: (2025)
by: Harel, Itamar, et al.
Published: (2025)
SKIM: Any-bit Quantization Pushing The Limits of Post-Training Quantization
by: Bai, Runsheng, et al.
Published: (2024)
by: Bai, Runsheng, et al.
Published: (2024)
A Family of Kernelized Matrix Costs for Multiple-Output Mixture Neural Networks
by: Hu, Bo, et al.
Published: (2025)
by: Hu, Bo, et al.
Published: (2025)
BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate
by: Mazza, Arnon, et al.
Published: (2026)
by: Mazza, Arnon, et al.
Published: (2026)
Learning Rate Scheduling with Matrix Factorization for Private Training
by: Kalinin, Nikita P., et al.
Published: (2025)
by: Kalinin, Nikita P., et al.
Published: (2025)
Ultra-Quantisation: Efficient Embedding Search via 1.58-bit Encodings
by: Connor, Richard, et al.
Published: (2025)
by: Connor, Richard, et al.
Published: (2025)
Similar Items
-
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025) -
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
by: Chmiel, Brian, et al.
Published: (2022) -
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024) -
EXAQ: Exponent Aware Quantization For LLMs Acceleration
by: Shkolnik, Moran, et al.
Published: (2024) -
Normalized Architectures are Natively 4-Bit
by: Fishman, Maxim, et al.
Published: (2026)