StatQAT: Statistical Quantizer Optimization for Deep Networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Aktukmak, Mehmet, Huang, Daniel, Ding, Ke |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WPMixer: Efficient Multi-Resolution Mixing for Long-Term Time Series Forecasting
von: Murad, Md Mahmuddun Nabi, et al.
Veröffentlicht: (2024)
von: Murad, Md Mahmuddun Nabi, et al.
Veröffentlicht: (2024)
EfQAT: An Efficient Framework for Quantization-Aware Training
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)
BCJR-QAT: A Differentiable Relaxation of Trellis-Coded Weight Quantization
von: Iyengar, Venugopalan
Veröffentlicht: (2026)
von: Iyengar, Venugopalan
Veröffentlicht: (2026)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2026)
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
von: Gernigon, Cédric, et al.
Veröffentlicht: (2024)
von: Gernigon, Cédric, et al.
Veröffentlicht: (2024)
DivQAT: Enhancing Robustness of Quantized Convolutional Neural Networks against Model Extraction Attacks
von: Khaled, Kacem, et al.
Veröffentlicht: (2025)
von: Khaled, Kacem, et al.
Veröffentlicht: (2025)
StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths
von: Chen, Tianyi, et al.
Veröffentlicht: (2026)
von: Chen, Tianyi, et al.
Veröffentlicht: (2026)
MF-QAT: Multi-Format Quantization-Aware Training for Elastic Inference
von: Xu, Zifei, et al.
Veröffentlicht: (2026)
von: Xu, Zifei, et al.
Veröffentlicht: (2026)
Optimizing Large Language Models through Quantization: A Comparative Analysis of PTQ and QAT Techniques
von: Hasan, Jahid
Veröffentlicht: (2024)
von: Hasan, Jahid
Veröffentlicht: (2024)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
1-Bit Wonder: Improving QAT Performance in the Low-Bit Regime through K-Means Quantization
von: Maskey, Sohir, et al.
Veröffentlicht: (2026)
von: Maskey, Sohir, et al.
Veröffentlicht: (2026)
Efficient VQ-QAT and Mixed Vector/Linear quantized Neural Networks
von: Gou, Terry, et al.
Veröffentlicht: (2026)
von: Gou, Terry, et al.
Veröffentlicht: (2026)
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2025)
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2025)
RealStats: A Rigorous Real-Only Statistical Framework for Fake Image Detection
von: Zisman, Haim, et al.
Veröffentlicht: (2026)
von: Zisman, Haim, et al.
Veröffentlicht: (2026)
Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
von: Xu, Binxing, et al.
Veröffentlicht: (2026)
von: Xu, Binxing, et al.
Veröffentlicht: (2026)
Rethinking Deep Learning: Propagating Information in Neural Networks without Backpropagation and Statistical Optimization
von: Itoh, Kei
Veröffentlicht: (2024)
von: Itoh, Kei
Veröffentlicht: (2024)
Quantization-Aware Regularizers for Deep Neural Networks Compression
von: Malchiodi, Dario, et al.
Veröffentlicht: (2026)
von: Malchiodi, Dario, et al.
Veröffentlicht: (2026)
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
Quantization through Piecewise-Affine Regularization: Optimization and Statistical Guarantees
von: Ma, Jianhao, et al.
Veröffentlicht: (2025)
von: Ma, Jianhao, et al.
Veröffentlicht: (2025)
Statistically-Lossless Quantization of Large Language Models
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
RefineStat: Efficient Exploration for Probabilistic Program Synthesis
von: Kanda, Madhav, et al.
Veröffentlicht: (2025)
von: Kanda, Madhav, et al.
Veröffentlicht: (2025)
Forget by Uncertainty: Orthogonal Entropy Unlearning for Quantized Neural Networks
von: Zhang, Tian, et al.
Veröffentlicht: (2026)
von: Zhang, Tian, et al.
Veröffentlicht: (2026)
SigmaMedStat: Temporal Signal Modeling for ICU False Alarm Reduction
von: Ramachandran, Arunkumar
Veröffentlicht: (2026)
von: Ramachandran, Arunkumar
Veröffentlicht: (2026)
Effective Quantization of Muon Optimizer States
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
von: Behnia, Tina, et al.
Veröffentlicht: (2025)
von: Behnia, Tina, et al.
Veröffentlicht: (2025)
Detection of Odor Presence via Deep Neural Networks
von: Hassanloo, Matin, et al.
Veröffentlicht: (2025)
von: Hassanloo, Matin, et al.
Veröffentlicht: (2025)
Binary Feature Mask Optimization for Feature Selection
von: Lorasdagi, Mehmet E., et al.
Veröffentlicht: (2024)
von: Lorasdagi, Mehmet E., et al.
Veröffentlicht: (2024)
Mini-batch Estimation for Deep Cox Models: Statistical Foundations and Practical Guidance
von: Zeng, Lang, et al.
Veröffentlicht: (2024)
von: Zeng, Lang, et al.
Veröffentlicht: (2024)
Joint Pruning and Channel-wise Mixed-Precision Quantization for Efficient Deep Neural Networks
von: Motetti, Beatrice Alessandra, et al.
Veröffentlicht: (2024)
von: Motetti, Beatrice Alessandra, et al.
Veröffentlicht: (2024)
On the Universal Statistical Consistency of Expansive Hyperbolic Deep Convolutional Neural Networks
von: Ghosh, Sagar, et al.
Veröffentlicht: (2024)
von: Ghosh, Sagar, et al.
Veröffentlicht: (2024)
DQA: An Efficient Method for Deep Quantization of Deep Neural Network Activations
von: Hu, Wenhao, et al.
Veröffentlicht: (2024)
von: Hu, Wenhao, et al.
Veröffentlicht: (2024)
Wahkon: A Statistically Principled Deep RKHS Superposition Network
von: Chen, Yongkai, et al.
Veröffentlicht: (2026)
von: Chen, Yongkai, et al.
Veröffentlicht: (2026)
Low-bit Model Quantization for Deep Neural Networks: A Survey
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
InfoQ: Mixed-Precision Quantization via Global Information Flow
von: Akbulut, Mehmet Emre, et al.
Veröffentlicht: (2025)
von: Akbulut, Mehmet Emre, et al.
Veröffentlicht: (2025)
Critical Organization of Deep Neural Networks, and p-Adic Statistical Field Theories
von: Zúñiga-Galindo, W. A.
Veröffentlicht: (2026)
von: Zúñiga-Galindo, W. A.
Veröffentlicht: (2026)
The Binary Quantized Neural Network for Dense Prediction via Specially Designed Upsampling and Attention
von: Ding, Xingyu, et al.
Veröffentlicht: (2024)
von: Ding, Xingyu, et al.
Veröffentlicht: (2024)
Rethinking Post-Training Quantization: Introducing a Statistical Pre-Calibration Approach
von: Ghaffari, Alireza, et al.
Veröffentlicht: (2025)
von: Ghaffari, Alireza, et al.
Veröffentlicht: (2025)
On-Device Training of Fully Quantized Deep Neural Networks on Cortex-M Microcontrollers
von: Deutel, Mark, et al.
Veröffentlicht: (2024)
von: Deutel, Mark, et al.
Veröffentlicht: (2024)
Exploring Quantization and Mapping Synergy in Hardware-Aware Deep Neural Network Accelerators
von: Klhufek, Jan, et al.
Veröffentlicht: (2024)
von: Klhufek, Jan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
WPMixer: Efficient Multi-Resolution Mixing for Long-Term Time Series Forecasting
von: Murad, Md Mahmuddun Nabi, et al.
Veröffentlicht: (2024) -
EfQAT: An Efficient Framework for Quantization-Aware Training
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024) -
BCJR-QAT: A Differentiable Relaxation of Trellis-Coded Weight Quantization
von: Iyengar, Venugopalan
Veröffentlicht: (2026) -
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2026) -
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
von: Gernigon, Cédric, et al.
Veröffentlicht: (2024)