Saved in:
| Main Authors: | Bablani, Deepika, Mckinstry, Jeffrey L., Esser, Steven K., Appuswamy, Rathinakumar, Modha, Dharmendra S. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2301.13330 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SiLQ: Simple Large Language Model Quantization-Aware Training
by: Esser, Steven K., et al.
Published: (2025)
by: Esser, Steven K., et al.
Published: (2025)
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference
by: Gafni, Tomer, et al.
Published: (2025)
by: Gafni, Tomer, et al.
Published: (2025)
Adaptive Distribution-aware Quantization for Mixed-Precision Neural Networks
by: Jia, Shaohang, et al.
Published: (2025)
by: Jia, Shaohang, et al.
Published: (2025)
Probability of super-regular matrices and MDS codes over finite fields
by: Appuswamy, Rathinakumar, et al.
Published: (2026)
by: Appuswamy, Rathinakumar, et al.
Published: (2026)
Value-Driven Mixed-Precision Quantization for Patch-Based Inference on Microcontrollers
by: Tao, Wei, et al.
Published: (2024)
by: Tao, Wei, et al.
Published: (2024)
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
by: DeBole, Michael V., et al.
Published: (2025)
by: DeBole, Michael V., et al.
Published: (2025)
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
by: Tao, Wei, et al.
Published: (2025)
by: Tao, Wei, et al.
Published: (2025)
6Bit-Diffusion: Inference-Time Mixed-Precision Quantization for Video Diffusion Models
by: Su, Rundong, et al.
Published: (2026)
by: Su, Rundong, et al.
Published: (2026)
Mix-QSAM: Mixed-Precision Quantization of the Segment Anything Model
by: Ranjan, Navin, et al.
Published: (2025)
by: Ranjan, Navin, et al.
Published: (2025)
Mix-QViT: Mixed-Precision Vision Transformer Quantization Driven by Layer Importance and Quantization Sensitivity
by: Ranjan, Navin, et al.
Published: (2025)
by: Ranjan, Navin, et al.
Published: (2025)
Efficient Mixed Precision Quantization in Graph Neural Networks
by: Moustafa, Samir, et al.
Published: (2025)
by: Moustafa, Samir, et al.
Published: (2025)
DQA: An Efficient Method for Deep Quantization of Deep Neural Network Activations
by: Hu, Wenhao, et al.
Published: (2024)
by: Hu, Wenhao, et al.
Published: (2024)
Precision Neural Network Quantization via Learnable Adaptive Modules
by: Zhou, Wenqiang, et al.
Published: (2025)
by: Zhou, Wenqiang, et al.
Published: (2025)
MPQ-Diff: Mixed Precision Quantization for Diffusion Models
by: Maruzzelli, Rocco Manz, et al.
Published: (2024)
by: Maruzzelli, Rocco Manz, et al.
Published: (2024)
MP-DPD: Low-Complexity Mixed-Precision Neural Networks for Energy-Efficient Digital Predistortion of Wideband Power Amplifiers
by: Wu, Yizhuo, et al.
Published: (2024)
by: Wu, Yizhuo, et al.
Published: (2024)
Real-Time Spacecraft Pose Estimation Using Mixed-Precision Quantized Neural Network on COTS Reconfigurable MPSoC
by: Posso, Julien, et al.
Published: (2024)
by: Posso, Julien, et al.
Published: (2024)
MetaMix: Meta-state Precision Searcher for Mixed-precision Activation Quantization
by: Kim, Han-Byul, et al.
Published: (2023)
by: Kim, Han-Byul, et al.
Published: (2023)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
by: Tao, Wei, et al.
Published: (2026)
by: Tao, Wei, et al.
Published: (2026)
MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization
by: Zhao, Tianchen, et al.
Published: (2024)
by: Zhao, Tianchen, et al.
Published: (2024)
MPTQ-ViT: Mixed-Precision Post-Training Quantization for Vision Transformer
by: Tai, Yu-Shan, et al.
Published: (2024)
by: Tai, Yu-Shan, et al.
Published: (2024)
MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models
by: Feng, Weilun, et al.
Published: (2024)
by: Feng, Weilun, et al.
Published: (2024)
Smaller, Faster, Cheaper: Architectural Designs for Efficient Machine Learning
by: Walton, Steven
Published: (2025)
by: Walton, Steven
Published: (2025)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
by: Xu, Haoning, et al.
Published: (2025)
by: Xu, Haoning, et al.
Published: (2025)
DynaQuant: Dynamic Mixed-Precision Quantization for Learned Image Compression
by: Bao, Youneng, et al.
Published: (2025)
by: Bao, Youneng, et al.
Published: (2025)
MixA-Q: Revisiting Activation Sparsity for Vision Transformers from a Mixed-Precision Quantization Perspective
by: Wang, Weitian, et al.
Published: (2025)
by: Wang, Weitian, et al.
Published: (2025)
Data-Free Quantization via Mixed-Precision Compensation without Fine-Tuning
by: Chen, Jun, et al.
Published: (2023)
by: Chen, Jun, et al.
Published: (2023)
FLIQS: One-Shot Mixed-Precision Floating-Point and Integer Quantization Search
by: Dotzel, Jordan, et al.
Published: (2023)
by: Dotzel, Jordan, et al.
Published: (2023)
FairQuant: Fairness-Aware Mixed-Precision Quantization for Medical Image Classification
by: Woergaard, Thomas, et al.
Published: (2026)
by: Woergaard, Thomas, et al.
Published: (2026)
LampQ: Towards Accurate Layer-wise Mixed Precision Quantization for Vision Transformers
by: Kim, Minjun, et al.
Published: (2025)
by: Kim, Minjun, et al.
Published: (2025)
Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
by: Bagrov, Natan, et al.
Published: (2025)
by: Bagrov, Natan, et al.
Published: (2025)
ARQ: A Mixed-Precision Quantization Framework for Accurate and Certifiably Robust DNNs
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
CompGS: Smaller and Faster Gaussian Splatting with Vector Quantization
by: Navaneet, KL, et al.
Published: (2023)
by: Navaneet, KL, et al.
Published: (2023)
Restoring Neural Network Plasticity for Faster Transfer Learning
by: Coetzer, Xander, et al.
Published: (2026)
by: Coetzer, Xander, et al.
Published: (2026)
Algebraic Representations for Faster Predictions in Convolutional Neural Networks
by: Joyce, Johnny, et al.
Published: (2024)
by: Joyce, Johnny, et al.
Published: (2024)
LRP-QViT: Mixed-Precision Vision Transformer Quantization via Layer-wise Relevance Propagation
by: Ranjan, Navin, et al.
Published: (2024)
by: Ranjan, Navin, et al.
Published: (2024)
Joint Pruning and Channel-wise Mixed-Precision Quantization for Efficient Deep Neural Networks
by: Motetti, Beatrice Alessandra, et al.
Published: (2024)
by: Motetti, Beatrice Alessandra, et al.
Published: (2024)
TreeQ: Pushing the Quantization Boundary of Diffusion Transformer via Tree-Structured Mixed-Precision Search
by: Yang, Kaicheng, et al.
Published: (2025)
by: Yang, Kaicheng, et al.
Published: (2025)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
Tiny-PULP-Dronets: Squeezing Neural Networks for Faster and Lighter Inference on Multi-Tasking Autonomous Nano-Drones
by: Lamberti, Lorenzo, et al.
Published: (2024)
by: Lamberti, Lorenzo, et al.
Published: (2024)
Efficiently Training A Flat Neural Network Before It has been Quantizated
by: Xia, Peng, et al.
Published: (2025)
by: Xia, Peng, et al.
Published: (2025)
Similar Items
-
SiLQ: Simple Large Language Model Quantization-Aware Training
by: Esser, Steven K., et al.
Published: (2025) -
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference
by: Gafni, Tomer, et al.
Published: (2025) -
Adaptive Distribution-aware Quantization for Mixed-Precision Neural Networks
by: Jia, Shaohang, et al.
Published: (2025) -
Probability of super-regular matrices and MDS codes over finite fields
by: Appuswamy, Rathinakumar, et al.
Published: (2026) -
Value-Driven Mixed-Precision Quantization for Patch-Based Inference on Microcontrollers
by: Tao, Wei, et al.
Published: (2024)