Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Gafni, Tomer, Karnieli, Asaf, Hanani, Yair |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dual Precision Deep Neural Network
by: Park, Jae Hyun, et al.
Published: (2020)
by: Park, Jae Hyun, et al.
Published: (2020)
Efficient and Effective Methods for Mixed Precision Neural Network Quantization for Faster, Energy-efficient Inference
by: Bablani, Deepika, et al.
Published: (2023)
by: Bablani, Deepika, et al.
Published: (2023)
Fine-tuning Quantized Neural Networks with Zeroth-order Optimization
by: Shang, Sifeng, et al.
Published: (2025)
by: Shang, Sifeng, et al.
Published: (2025)
DQA: An Efficient Method for Deep Quantization of Deep Neural Network Activations
by: Hu, Wenhao, et al.
Published: (2024)
by: Hu, Wenhao, et al.
Published: (2024)
IM-Unpack: Training and Inference with Arbitrarily Low Precision Integers
by: Zeng, Zhanpeng, et al.
Published: (2024)
by: Zeng, Zhanpeng, et al.
Published: (2024)
DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
by: Picón, Ginés Carreto, et al.
Published: (2025)
by: Picón, Ginés Carreto, et al.
Published: (2025)
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
by: Miranda, Imanol, et al.
Published: (2026)
by: Miranda, Imanol, et al.
Published: (2026)
First-Order Error Matters: Accurate Compensation for Quantized Large Language Models
by: Zheng, Xingyu, et al.
Published: (2025)
by: Zheng, Xingyu, et al.
Published: (2025)
Till the Layers Collapse: Compressing a Deep Neural Network through the Lenses of Batch Normalization Layers
by: Liao, Zhu, et al.
Published: (2024)
by: Liao, Zhu, et al.
Published: (2024)
Growing Efficient Accurate and Robust Neural Networks on the Edge
by: Sundaresha, Vignesh, et al.
Published: (2024)
by: Sundaresha, Vignesh, et al.
Published: (2024)
A Primal-Dual Framework for Transformers and Neural Networks
by: Nguyen, Tan M., et al.
Published: (2024)
by: Nguyen, Tan M., et al.
Published: (2024)
Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization
by: Wang, YiFeng, et al.
Published: (2026)
by: Wang, YiFeng, et al.
Published: (2026)
Value-Driven Mixed-Precision Quantization for Patch-Based Inference on Microcontrollers
by: Tao, Wei, et al.
Published: (2024)
by: Tao, Wei, et al.
Published: (2024)
Quantized and Interpretable Learning Scheme for Deep Neural Networks in Classification Task
by: Maleki, Alireza, et al.
Published: (2024)
by: Maleki, Alireza, et al.
Published: (2024)
Recurrent Neural Networks for Still Images
by: Dmitri, et al.
Published: (2024)
by: Dmitri, et al.
Published: (2024)
Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies
by: Pathak, Surendra, et al.
Published: (2026)
by: Pathak, Surendra, et al.
Published: (2026)
BayesQ: Uncertainty-Guided Bayesian Quantization
by: Lamaakal, Ismail, et al.
Published: (2025)
by: Lamaakal, Ismail, et al.
Published: (2025)
Scaling Up Quantization-Aware Neural Architecture Search for Efficient Deep Learning on the Edge
by: Lu, Yao, et al.
Published: (2024)
by: Lu, Yao, et al.
Published: (2024)
Robust Training of Neural Networks at Arbitrary Precision and Sparsity
by: Ye, Chengxi, et al.
Published: (2024)
by: Ye, Chengxi, et al.
Published: (2024)
Enhancing Post-Training Quantization via Future Activation Awareness
by: Lv, Zheqi, et al.
Published: (2026)
by: Lv, Zheqi, et al.
Published: (2026)
Transformer-VQ: Linear-Time Transformers via Vector Quantization
by: Lingle, Lucas D.
Published: (2023)
by: Lingle, Lucas D.
Published: (2023)
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
Benchmarking Graph Neural Networks for Document Layout Analysis in Public Affairs
by: Lopez-Duran, Miguel, et al.
Published: (2025)
by: Lopez-Duran, Miguel, et al.
Published: (2025)
VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models
by: Hou, Haowen, et al.
Published: (2024)
by: Hou, Haowen, et al.
Published: (2024)
Saliency Assisted Quantization for Neural Networks
by: Rezabeyk, Elmira Mousa, et al.
Published: (2024)
by: Rezabeyk, Elmira Mousa, et al.
Published: (2024)
Coherence Awareness in Diffractive Neural Networks
by: Kleiner, Matan, et al.
Published: (2024)
by: Kleiner, Matan, et al.
Published: (2024)
Expansion Quantization Network: An Efficient Micro-emotion Annotation and Detection Framework
by: Zhou, Jingyi, et al.
Published: (2024)
by: Zhou, Jingyi, et al.
Published: (2024)
VLCD: Vision-Language Contrastive Distillation for Accurate and Efficient Automatic Placenta Analysis
by: Mehta, Manas, et al.
Published: (2025)
by: Mehta, Manas, et al.
Published: (2025)
TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
by: Liang, Guang, et al.
Published: (2025)
by: Liang, Guang, et al.
Published: (2025)
ARQ: A Mixed-Precision Quantization Framework for Accurate and Certifiably Robust DNNs
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
FisherMask: Enhancing Neural Network Labeling Efficiency in Image Classification Using Fisher Information
by: Gul, Shreen, et al.
Published: (2024)
by: Gul, Shreen, et al.
Published: (2024)
Dual-Process Image Generation
by: Luo, Grace, et al.
Published: (2025)
by: Luo, Grace, et al.
Published: (2025)
TMPQ-DM: Joint Timestep Reduction and Quantization Precision Selection for Efficient Diffusion Models
by: Sun, Haojun, et al.
Published: (2024)
by: Sun, Haojun, et al.
Published: (2024)
MoENAS: Mixture-of-Expert based Neural Architecture Search for jointly Accurate, Fair, and Robust Edge Deep Neural Networks
by: Mecharbat, Lotfi Abdelkrim, et al.
Published: (2025)
by: Mecharbat, Lotfi Abdelkrim, et al.
Published: (2025)
MatFormer: Nested Transformer for Elastic Inference
by: Devvrit, et al.
Published: (2023)
by: Devvrit, et al.
Published: (2023)
Q-SENN: Quantized Self-Explaining Neural Networks
by: Norrenbrock, Thomas, et al.
Published: (2023)
by: Norrenbrock, Thomas, et al.
Published: (2023)
No Training Wheels: Steering Vectors for Bias Correction at Inference Time
by: Gupta, Aviral, et al.
Published: (2025)
by: Gupta, Aviral, et al.
Published: (2025)
A General Framework for Inference-time Scaling and Steering of Diffusion Models
by: Singhal, Raghav, et al.
Published: (2025)
by: Singhal, Raghav, et al.
Published: (2025)
Multimodal Adaptive Inference for Document Image Classification with Anytime Early Exiting
by: Hamed, Omar, et al.
Published: (2024)
by: Hamed, Omar, et al.
Published: (2024)
ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time
by: Ding, Yi, et al.
Published: (2024)
by: Ding, Yi, et al.
Published: (2024)
Similar Items
-
Dual Precision Deep Neural Network
by: Park, Jae Hyun, et al.
Published: (2020) -
Efficient and Effective Methods for Mixed Precision Neural Network Quantization for Faster, Energy-efficient Inference
by: Bablani, Deepika, et al.
Published: (2023) -
Fine-tuning Quantized Neural Networks with Zeroth-order Optimization
by: Shang, Sifeng, et al.
Published: (2025) -
DQA: An Efficient Method for Deep Quantization of Deep Neural Network Activations
by: Hu, Wenhao, et al.
Published: (2024) -
IM-Unpack: Training and Inference with Arbitrarily Low Precision Integers
by: Zeng, Zhanpeng, et al.
Published: (2024)