Improving Quantization-aware Training of Low-Precision Network via Block Replacement on Full-Precision Counterpart
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Chengting, Yang, Shu, Zhang, Fengzhao, Ma, Hanzhi, Wang, Aili, Li, Er-Ping |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Go beyond End-to-End Training: Boosting Greedy Local Learning with Context Supply
by: Yu, Chengting, et al.
Published: (2023)
by: Yu, Chengting, et al.
Published: (2023)
Efficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise Replacement
by: Yang, Shu, et al.
Published: (2025)
by: Yang, Shu, et al.
Published: (2025)
Decoupling Dark Knowledge via Block-wise Logit Distillation for Feature-level Alignment
by: Yu, Chengting, et al.
Published: (2024)
by: Yu, Chengting, et al.
Published: (2024)
SDiT: Spiking Diffusion Model with Transformer
by: Yang, Shu, et al.
Published: (2024)
by: Yang, Shu, et al.
Published: (2024)
Efficient Logit-based Knowledge Distillation of Deep Spiking Neural Networks for Full-Range Timestep Deployment
by: Yu, Chengting, et al.
Published: (2025)
by: Yu, Chengting, et al.
Published: (2025)
Advancing Training Efficiency of Deep Spiking Neural Networks through Rate-based Backpropagation
by: Yu, Chengting, et al.
Published: (2024)
by: Yu, Chengting, et al.
Published: (2024)
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
by: Zhou, Cyrus, et al.
Published: (2023)
by: Zhou, Cyrus, et al.
Published: (2023)
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)
by: Nikdan, Mahdi, et al.
Published: (2026)
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
by: Xu, Zifei, et al.
Published: (2024)
by: Xu, Zifei, et al.
Published: (2024)
True 4-Bit Quantized Convolutional Neural Network Training on CPU: Achieving Full-Precision Parity
by: Tathe, Shivnath
Published: (2026)
by: Tathe, Shivnath
Published: (2026)
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
by: Guan, Ziyi, et al.
Published: (2024)
by: Guan, Ziyi, et al.
Published: (2024)
OMPQ: Orthogonal Mixed Precision Quantization
by: Ma, Yuexiao, et al.
Published: (2021)
by: Ma, Yuexiao, et al.
Published: (2021)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
by: Federici, Marco, et al.
Published: (2025)
by: Federici, Marco, et al.
Published: (2025)
BAPS: A Fine-Grained Low-Precision Scheme for Softmax in Attention via Block-Aware Precision reScaling
by: Ye, Zisheng, et al.
Published: (2026)
by: Ye, Zisheng, et al.
Published: (2026)
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
by: Liu, Wenyuan, et al.
Published: (2024)
by: Liu, Wenyuan, et al.
Published: (2024)
Efficient Mixed Precision Quantization in Graph Neural Networks
by: Moustafa, Samir, et al.
Published: (2025)
by: Moustafa, Samir, et al.
Published: (2025)
CPT: Efficient Deep Neural Network Training via Cyclic Precision
by: Fu, Yonggan, et al.
Published: (2021)
by: Fu, Yonggan, et al.
Published: (2021)
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
by: Huang, Xijie, et al.
Published: (2023)
by: Huang, Xijie, et al.
Published: (2023)
Where and How to Enhance: Discovering Bit-Width Contribution for Mixed Precision Quantization
by: Kang, Haidong, et al.
Published: (2025)
by: Kang, Haidong, et al.
Published: (2025)
Log-Normal Multiplicative Dynamics for Stable Low-Precision Training of Large Networks
by: Nishida, Keigo, et al.
Published: (2025)
by: Nishida, Keigo, et al.
Published: (2025)
Collage: Light-Weight Low-Precision Strategy for LLM Training
by: Yu, Tao, et al.
Published: (2024)
by: Yu, Tao, et al.
Published: (2024)
SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights
by: Müller, Lorenz K., et al.
Published: (2025)
by: Müller, Lorenz K., et al.
Published: (2025)
LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
by: Mirzaei, Amir Reza, et al.
Published: (2025)
by: Mirzaei, Amir Reza, et al.
Published: (2025)
DNN Memory Footprint Reduction via Post-Training Intra-Layer Multi-Precision Quantization
by: Ghavami, Behnam, et al.
Published: (2024)
by: Ghavami, Behnam, et al.
Published: (2024)
CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts
by: Yin, Xiangyang, et al.
Published: (2026)
by: Yin, Xiangyang, et al.
Published: (2026)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
by: Lee, Seoungsub, et al.
Published: (2026)
by: Lee, Seoungsub, et al.
Published: (2026)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
by: Dadgarnia, Alireza, et al.
Published: (2026)
by: Dadgarnia, Alireza, et al.
Published: (2026)
InfoQ: Mixed-Precision Quantization via Global Information Flow
by: Akbulut, Mehmet Emre, et al.
Published: (2025)
by: Akbulut, Mehmet Emre, et al.
Published: (2025)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
Adaptive Low-Precision Training for Embeddings in Click-Through Rate Prediction
by: Li, Shiwei, et al.
Published: (2022)
by: Li, Shiwei, et al.
Published: (2022)
AutoQRA: Joint Optimization of Mixed-Precision Quantization and Low-rank Adapters for Efficient LLM Fine-Tuning
by: Zhou, Changhai, et al.
Published: (2026)
by: Zhou, Changhai, et al.
Published: (2026)
AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation
by: Kim, Seonggon, et al.
Published: (2026)
by: Kim, Seonggon, et al.
Published: (2026)
LoPRo: Enhancing Low-Rank Quantization via Permuted Block-Wise Rotation
by: Gu, Hongyaoxing, et al.
Published: (2026)
by: Gu, Hongyaoxing, et al.
Published: (2026)
Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
by: Hao, Zhiwei, et al.
Published: (2025)
by: Hao, Zhiwei, et al.
Published: (2025)
TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction
by: Li, Yuhang, et al.
Published: (2024)
by: Li, Yuhang, et al.
Published: (2024)
AMED: Automatic Mixed-Precision Quantization for Edge Devices
by: Kimhi, Moshe, et al.
Published: (2022)
by: Kimhi, Moshe, et al.
Published: (2022)
MoPEQ: Mixture of Mixed Precision Quantized Experts
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
CoDeQ: End-to-End Joint Model Compression with Dead-Zone Quantizer for High-Sparsity and Low-Precision Networks
by: Wenshøj, Jonathan, et al.
Published: (2025)
by: Wenshøj, Jonathan, et al.
Published: (2025)
Joint Pruning and Channel-wise Mixed-Precision Quantization for Efficient Deep Neural Networks
by: Motetti, Beatrice Alessandra, et al.
Published: (2024)
by: Motetti, Beatrice Alessandra, et al.
Published: (2024)
DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network Inference
by: Zhou, Jiajun, et al.
Published: (2023)
by: Zhou, Jiajun, et al.
Published: (2023)
Similar Items
-
Go beyond End-to-End Training: Boosting Greedy Local Learning with Context Supply
by: Yu, Chengting, et al.
Published: (2023) -
Efficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise Replacement
by: Yang, Shu, et al.
Published: (2025) -
Decoupling Dark Knowledge via Block-wise Logit Distillation for Feature-level Alignment
by: Yu, Chengting, et al.
Published: (2024) -
SDiT: Spiking Diffusion Model with Transformer
by: Yang, Shu, et al.
Published: (2024) -
Efficient Logit-based Knowledge Distillation of Deep Spiking Neural Networks for Full-Range Timestep Deployment
by: Yu, Chengting, et al.
Published: (2025)