True 4-Bit Quantized Convolutional Neural Network Training on CPU: Achieving Full-Precision Parity
Fuente:
arXiv
Guardado en:
| Autor principal: | Tathe, Shivnath |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LACE: Loss-Adaptive Capacity Expansion for Continual Learning
por: Tathe, Shivnath
Publicado: (2026)
por: Tathe, Shivnath
Publicado: (2026)
DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network Inference
por: Zhou, Jiajun, et al.
Publicado: (2023)
por: Zhou, Jiajun, et al.
Publicado: (2023)
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
por: Zhou, Cyrus, et al.
Publicado: (2023)
por: Zhou, Cyrus, et al.
Publicado: (2023)
Improving Quantization-aware Training of Low-Precision Network via Block Replacement on Full-Precision Counterpart
por: Yu, Chengting, et al.
Publicado: (2024)
por: Yu, Chengting, et al.
Publicado: (2024)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
por: Zhang, Peiyuan, et al.
Publicado: (2026)
por: Zhang, Peiyuan, et al.
Publicado: (2026)
ECO: Quantized Training without Full-Precision Master Weights
por: Nikdan, Mahdi, et al.
Publicado: (2026)
por: Nikdan, Mahdi, et al.
Publicado: (2026)
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
por: Peng, Yanxin, et al.
Publicado: (2025)
por: Peng, Yanxin, et al.
Publicado: (2025)
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
por: Huang, Xijie, et al.
Publicado: (2023)
por: Huang, Xijie, et al.
Publicado: (2023)
Efficient Mixed Precision Quantization in Graph Neural Networks
por: Moustafa, Samir, et al.
Publicado: (2025)
por: Moustafa, Samir, et al.
Publicado: (2025)
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
por: Song, Jaewoo, et al.
Publicado: (2025)
por: Song, Jaewoo, et al.
Publicado: (2025)
Multiscale Training of Convolutional Neural Networks
por: Ahamed, Shadab, et al.
Publicado: (2025)
por: Ahamed, Shadab, et al.
Publicado: (2025)
Where and How to Enhance: Discovering Bit-Width Contribution for Mixed Precision Quantization
por: Kang, Haidong, et al.
Publicado: (2025)
por: Kang, Haidong, et al.
Publicado: (2025)
LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
por: Mirzaei, Amir Reza, et al.
Publicado: (2025)
por: Mirzaei, Amir Reza, et al.
Publicado: (2025)
TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
por: Kim, Jinhee, et al.
Publicado: (2025)
por: Kim, Jinhee, et al.
Publicado: (2025)
Verification of Bit-Flip Attacks against Quantized Neural Networks
por: Zhang, Yedi, et al.
Publicado: (2025)
por: Zhang, Yedi, et al.
Publicado: (2025)
Quantized Convolutional Neural Networks Through the Lens of Partial Differential Equations
por: Ben-Yair, Ido, et al.
Publicado: (2021)
por: Ben-Yair, Ido, et al.
Publicado: (2021)
CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs
por: Wang, Haoyu, et al.
Publicado: (2024)
por: Wang, Haoyu, et al.
Publicado: (2024)
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
por: Gernigon, Cédric, et al.
Publicado: (2024)
por: Gernigon, Cédric, et al.
Publicado: (2024)
Gradient-Free Training of Quantized Neural Networks
por: Cohen, Noa, et al.
Publicado: (2024)
por: Cohen, Noa, et al.
Publicado: (2024)
Hardness of Learning Fixed Parities with Neural Networks
por: Shoshani, Itamar, et al.
Publicado: (2025)
por: Shoshani, Itamar, et al.
Publicado: (2025)
Every Bit Counts: A Theoretical Study of Precision-Expressivity Tradeoffs in Quantized Transformers
por: Chakrabarti, Sayak, et al.
Publicado: (2026)
por: Chakrabarti, Sayak, et al.
Publicado: (2026)
Bits for Privacy: Evaluating Post-Training Quantization via Membership Inference
por: Zhang, Chenxiang, et al.
Publicado: (2025)
por: Zhang, Chenxiang, et al.
Publicado: (2025)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
por: Park, Jungwoo, et al.
Publicado: (2025)
por: Park, Jungwoo, et al.
Publicado: (2025)
HBVLA: Pushing 1-Bit Post-Training Quantization for Vision-Language-Action Models
por: Yan, Xin, et al.
Publicado: (2026)
por: Yan, Xin, et al.
Publicado: (2026)
TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction
por: Li, Yuhang, et al.
Publicado: (2024)
por: Li, Yuhang, et al.
Publicado: (2024)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
por: Ouyang, Xu, et al.
Publicado: (2024)
por: Ouyang, Xu, et al.
Publicado: (2024)
Outlier-Aware Training for Low-Bit Quantization of Structural Re-Parameterized Networks
por: Niu, Muqun, et al.
Publicado: (2024)
por: Niu, Muqun, et al.
Publicado: (2024)
NoProp: Training Neural Networks without Full Back-propagation or Full Forward-propagation
por: Li, Qinyu, et al.
Publicado: (2025)
por: Li, Qinyu, et al.
Publicado: (2025)
FAMES: Fast Approximate Multiplier Substitution for Mixed-Precision Quantized DNNs--Down to 2 Bits!
por: Ren, Yi, et al.
Publicado: (2024)
por: Ren, Yi, et al.
Publicado: (2024)
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
por: Jia, Jinda, et al.
Publicado: (2024)
por: Jia, Jinda, et al.
Publicado: (2024)
Joint Pruning and Channel-wise Mixed-Precision Quantization for Efficient Deep Neural Networks
por: Motetti, Beatrice Alessandra, et al.
Publicado: (2024)
por: Motetti, Beatrice Alessandra, et al.
Publicado: (2024)
Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate
por: Kim, Jaemin, et al.
Publicado: (2026)
por: Kim, Jaemin, et al.
Publicado: (2026)
MSQ: Memory-Efficient Bit Sparsification Quantization
por: Han, Seokho, et al.
Publicado: (2025)
por: Han, Seokho, et al.
Publicado: (2025)
One-Bit Quantization for Random Features Models
por: Akhtiamov, Danil, et al.
Publicado: (2025)
por: Akhtiamov, Danil, et al.
Publicado: (2025)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
por: Huang, Wei, et al.
Publicado: (2023)
por: Huang, Wei, et al.
Publicado: (2023)
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
por: Xu, Zifei, et al.
Publicado: (2024)
por: Xu, Zifei, et al.
Publicado: (2024)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
por: Wang, Dongwei, et al.
Publicado: (2026)
por: Wang, Dongwei, et al.
Publicado: (2026)
True Online TD-Replan(lambda) Achieving Planning through Replaying
por: Altahhan, Abdulrahman
Publicado: (2025)
por: Altahhan, Abdulrahman
Publicado: (2025)
Convolutional Neural Network Achieves Human-level Accuracy in Music Genre Classification
por: Dong, Mingwen
Publicado: (2018)
por: Dong, Mingwen
Publicado: (2018)
BitLogic: Training Framework for Gradient-Based FPGA-Native Neural Networks
por: Bührer, Simon, et al.
Publicado: (2026)
por: Bührer, Simon, et al.
Publicado: (2026)
Ejemplares similares
-
LACE: Loss-Adaptive Capacity Expansion for Continual Learning
por: Tathe, Shivnath
Publicado: (2026) -
DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network Inference
por: Zhou, Jiajun, et al.
Publicado: (2023) -
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
por: Zhou, Cyrus, et al.
Publicado: (2023) -
Improving Quantization-aware Training of Low-Precision Network via Block Replacement on Full-Precision Counterpart
por: Yu, Chengting, et al.
Publicado: (2024) -
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
por: Zhang, Peiyuan, et al.
Publicado: (2026)