DNN Memory Footprint Reduction via Post-Training Intra-Layer Multi-Precision Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Ghavami, Behnam, Kamjoo, Amin, Shannon, Lesley, Wilton, Steve |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ZOBNN: Zero-Overhead Dependable Design of Binary Neural Networks with Deliberately Quantized Parameters
by: Ghavami, Behnam, et al.
Published: (2024)
by: Ghavami, Behnam, et al.
Published: (2024)
Automatic High-quality Verilog Assertion Generation through Subtask-Focused Fine-Tuned LLMs and Iterative Prompting
by: Shahidzadeh, Mohammad, et al.
Published: (2024)
by: Shahidzadeh, Mohammad, et al.
Published: (2024)
Compressing Deep Neural Networks Using Explainable AI
by: Soroush, Kimia, et al.
Published: (2025)
by: Soroush, Kimia, et al.
Published: (2025)
A Semi Black-Box Adversarial Bit-Flip Attack with Limited DNN Model Information
by: Ghavami, Behnam, et al.
Published: (2024)
by: Ghavami, Behnam, et al.
Published: (2024)
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
by: Xie, Jianhang, et al.
Published: (2025)
by: Xie, Jianhang, et al.
Published: (2025)
MagR: Weight Magnitude Reduction for Enhancing Post-Training Quantization
by: Zhang, Aozhong, et al.
Published: (2024)
by: Zhang, Aozhong, et al.
Published: (2024)
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
by: Liu, Wenyuan, et al.
Published: (2024)
by: Liu, Wenyuan, et al.
Published: (2024)
Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
by: Xiao, He, et al.
Published: (2025)
by: Xiao, He, et al.
Published: (2025)
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
by: Guan, Ziyi, et al.
Published: (2024)
by: Guan, Ziyi, et al.
Published: (2024)
Pushing the Limits of Block Rotations in Post-Training Quantization
by: Sanjeet, Sai, et al.
Published: (2026)
by: Sanjeet, Sai, et al.
Published: (2026)
Beacon: Post-Training Quantization with Integrated Grid Selection
by: Zhang, Shihao, et al.
Published: (2025)
by: Zhang, Shihao, et al.
Published: (2025)
A Quantized VAE-MLP Botnet Detection Model: A Systematic Evaluation of Quantization-Aware Training and Post-Training Quantization Strategies
by: Wasswa, Hassan, et al.
Published: (2025)
by: Wasswa, Hassan, et al.
Published: (2025)
Post Training Quantization of Large Language Models with Microscaling Formats
by: Sharify, Sayeh, et al.
Published: (2024)
by: Sharify, Sayeh, et al.
Published: (2024)
Activation Sensitivity as a Unifying Principle for Post-Training Quantization
by: Xu, Bruce Changlong
Published: (2026)
by: Xu, Bruce Changlong
Published: (2026)
SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
by: Luo, Yingsong, et al.
Published: (2024)
by: Luo, Yingsong, et al.
Published: (2024)
Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
by: Kim, Junhan, et al.
Published: (2024)
by: Kim, Junhan, et al.
Published: (2024)
Achieving binary weight and activation for LLMs using Post-Training Quantization
by: Song, Siqing, et al.
Published: (2025)
by: Song, Siqing, et al.
Published: (2025)
PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
by: Xiao, He, et al.
Published: (2025)
by: Xiao, He, et al.
Published: (2025)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
by: Yu, Xiaoming, et al.
Published: (2026)
by: Yu, Xiaoming, et al.
Published: (2026)
Improving Quantization with Post-Training Model Expansion
by: Franco, Giuseppe, et al.
Published: (2025)
by: Franco, Giuseppe, et al.
Published: (2025)
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
by: Salimi, Moein, et al.
Published: (2026)
by: Salimi, Moein, et al.
Published: (2026)
NeFT: Negative Feedback Training to Improve Robustness of Compute-In-Memory DNN Accelerators
by: Qin, Yifan, et al.
Published: (2023)
by: Qin, Yifan, et al.
Published: (2023)
Quamba: A Post-Training Quantization Recipe for Selective State Space Models
by: Chiang, Hung-Yueh, et al.
Published: (2024)
by: Chiang, Hung-Yueh, et al.
Published: (2024)
BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs
by: Zhao, Zhixiong, et al.
Published: (2026)
by: Zhao, Zhixiong, et al.
Published: (2026)
Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training
by: Varshney, Ayush K., et al.
Published: (2026)
by: Varshney, Ayush K., et al.
Published: (2026)
Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training
by: Nepal, Aadim, et al.
Published: (2025)
by: Nepal, Aadim, et al.
Published: (2025)
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)
by: Nikdan, Mahdi, et al.
Published: (2026)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
by: Lee, Jung Hyun, et al.
Published: (2023)
by: Lee, Jung Hyun, et al.
Published: (2023)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
by: Yan, Xianglong, et al.
Published: (2026)
by: Yan, Xianglong, et al.
Published: (2026)
Astro: Activation-guided Structured Regularization for Outlier-Robust LLM Post-Training Quantization
by: Chen, Xi, et al.
Published: (2026)
by: Chen, Xi, et al.
Published: (2026)
RaanA: A Fast, Flexible, and Data-Efficient Post-Training Quantization Algorithm
by: Yang, Yongyi, et al.
Published: (2025)
by: Yang, Yongyi, et al.
Published: (2025)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
by: Zhao, Jiaqi, et al.
Published: (2025)
by: Zhao, Jiaqi, et al.
Published: (2025)
Post-Training Quantization of OpenPangu Models for Efficient Deployment on Atlas A2
by: Luo, Yilun, et al.
Published: (2025)
by: Luo, Yilun, et al.
Published: (2025)
A Unifying Post-Processing Framework for Multi-Objective Learn-to-Defer Problems
by: Charusaie, Mohammad-Amin, et al.
Published: (2024)
by: Charusaie, Mohammad-Amin, et al.
Published: (2024)
Accumulator-Aware Post-Training Quantization for Large Language Models
by: Colbert, Ian, et al.
Published: (2024)
by: Colbert, Ian, et al.
Published: (2024)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
Qronos: Correcting the Past by Shaping the Future... in Post-Training Quantization
by: Zhang, Shihao, et al.
Published: (2025)
by: Zhang, Shihao, et al.
Published: (2025)
MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
by: Zhao, Zhixiong, et al.
Published: (2026)
by: Zhao, Zhixiong, et al.
Published: (2026)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
by: Gadhikar, Advait, et al.
Published: (2025)
by: Gadhikar, Advait, et al.
Published: (2025)
Similar Items
-
ZOBNN: Zero-Overhead Dependable Design of Binary Neural Networks with Deliberately Quantized Parameters
by: Ghavami, Behnam, et al.
Published: (2024) -
Automatic High-quality Verilog Assertion Generation through Subtask-Focused Fine-Tuned LLMs and Iterative Prompting
by: Shahidzadeh, Mohammad, et al.
Published: (2024) -
Compressing Deep Neural Networks Using Explainable AI
by: Soroush, Kimia, et al.
Published: (2025) -
A Semi Black-Box Adversarial Bit-Flip Attack with Limited DNN Model Information
by: Ghavami, Behnam, et al.
Published: (2024) -
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
by: Xie, Jianhang, et al.
Published: (2025)