BCJR-QAT: A Differentiable Relaxation of Trellis-Coded Weight Quantization
Fuente:
arXiv
Salvato in:
| Autore principale: | Iyengar, Venugopalan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
StatQAT: Statistical Quantizer Optimization for Deep Networks
di: Aktukmak, Mehmet, et al.
Pubblicazione: (2026)
di: Aktukmak, Mehmet, et al.
Pubblicazione: (2026)
EfQAT: An Efficient Framework for Quantization-Aware Training
di: Ashkboos, Saleh, et al.
Pubblicazione: (2024)
di: Ashkboos, Saleh, et al.
Pubblicazione: (2024)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
di: Zhang, Peiyuan, et al.
Pubblicazione: (2026)
di: Zhang, Peiyuan, et al.
Pubblicazione: (2026)
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
di: Gernigon, Cédric, et al.
Pubblicazione: (2024)
di: Gernigon, Cédric, et al.
Pubblicazione: (2024)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
di: Ke, Wenjin, et al.
Pubblicazione: (2025)
di: Ke, Wenjin, et al.
Pubblicazione: (2025)
StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths
di: Chen, Tianyi, et al.
Pubblicazione: (2026)
di: Chen, Tianyi, et al.
Pubblicazione: (2026)
MF-QAT: Multi-Format Quantization-Aware Training for Elastic Inference
di: Xu, Zifei, et al.
Pubblicazione: (2026)
di: Xu, Zifei, et al.
Pubblicazione: (2026)
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
di: Chen, Mengzhao, et al.
Pubblicazione: (2024)
di: Chen, Mengzhao, et al.
Pubblicazione: (2024)
1-Bit Wonder: Improving QAT Performance in the Low-Bit Regime through K-Means Quantization
di: Maskey, Sohir, et al.
Pubblicazione: (2026)
di: Maskey, Sohir, et al.
Pubblicazione: (2026)
DivQAT: Enhancing Robustness of Quantized Convolutional Neural Networks against Model Extraction Attacks
di: Khaled, Kacem, et al.
Pubblicazione: (2025)
di: Khaled, Kacem, et al.
Pubblicazione: (2025)
Optimizing Large Language Models through Quantization: A Comparative Analysis of PTQ and QAT Techniques
di: Hasan, Jahid
Pubblicazione: (2024)
di: Hasan, Jahid
Pubblicazione: (2024)
Trellis: Learning to Compress Key-Value Memory in Attention Models
di: Karami, Mahdi, et al.
Pubblicazione: (2025)
di: Karami, Mahdi, et al.
Pubblicazione: (2025)
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
di: Lee, Jung Hyun, et al.
Pubblicazione: (2025)
di: Lee, Jung Hyun, et al.
Pubblicazione: (2025)
Efficient VQ-QAT and Mixed Vector/Linear quantized Neural Networks
di: Gou, Terry, et al.
Pubblicazione: (2026)
di: Gou, Terry, et al.
Pubblicazione: (2026)
Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
di: Xu, Binxing, et al.
Pubblicazione: (2026)
di: Xu, Binxing, et al.
Pubblicazione: (2026)
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
di: Zhang, Jintao, et al.
Pubblicazione: (2026)
di: Zhang, Jintao, et al.
Pubblicazione: (2026)
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
di: Xu, Yuzhuang, et al.
Pubblicazione: (2024)
di: Xu, Yuzhuang, et al.
Pubblicazione: (2024)
A Differentiable Bayesian Relaxation for Latent Partial-Order Inference
di: Li, Dongqing, et al.
Pubblicazione: (2026)
di: Li, Dongqing, et al.
Pubblicazione: (2026)
Trust via Reputation of Conviction
di: Iyengar, Aravind R.
Pubblicazione: (2026)
di: Iyengar, Aravind R.
Pubblicazione: (2026)
ADMM-Q: An Improved Hessian-based Weight Quantizer for Post-Training Quantization of Large Language Models
di: Lucas, Ryan, et al.
Pubblicazione: (2026)
di: Lucas, Ryan, et al.
Pubblicazione: (2026)
Effect of Weight Quantization on Learning Models by Typical Case Analysis
di: Kashiwamura, Shuhei, et al.
Pubblicazione: (2024)
di: Kashiwamura, Shuhei, et al.
Pubblicazione: (2024)
Restructuring Vector Quantization with the Rotation Trick
di: Fifty, Christopher, et al.
Pubblicazione: (2024)
di: Fifty, Christopher, et al.
Pubblicazione: (2024)
Model-Free Approximate Bayesian Learning for Large-Scale Conversion Funnel Optimization
di: Iyengar, Garud, et al.
Pubblicazione: (2024)
di: Iyengar, Garud, et al.
Pubblicazione: (2024)
Gaussian Weight Sampling for Scalable, Efficient and Stable Pseudo-Quantization Training
di: Ahn, Myeonghwan, et al.
Pubblicazione: (2025)
di: Ahn, Myeonghwan, et al.
Pubblicazione: (2025)
AWP: Activation-Aware Weight Pruning and Quantization with Projected Gradient Descent
di: Liu, Jing, et al.
Pubblicazione: (2025)
di: Liu, Jing, et al.
Pubblicazione: (2025)
Quantized Delta Weight Is Safety Keeper
di: Liu, Yule, et al.
Pubblicazione: (2024)
di: Liu, Yule, et al.
Pubblicazione: (2024)
Generalized Radius and Integrated Codebook Transforms for Differentiable Vector Quantization
di: You, Haochen, et al.
Pubblicazione: (2026)
di: You, Haochen, et al.
Pubblicazione: (2026)
The Cost of Learning under Multiple Change Points
di: Gafni, Tomer, et al.
Pubblicazione: (2026)
di: Gafni, Tomer, et al.
Pubblicazione: (2026)
Learning the Pareto Front Using Bootstrapped Observation Samples
di: Kim, Wonyoung, et al.
Pubblicazione: (2023)
di: Kim, Wonyoung, et al.
Pubblicazione: (2023)
A Giant-Step Baby-Step Classifier For Scalable and Real-Time Anomaly Detection In Industrial Control Systems and Water Treatment Systems
di: Venugopalan, Sarad, et al.
Pubblicazione: (2025)
di: Venugopalan, Sarad, et al.
Pubblicazione: (2025)
A Differentially Private Weighted Empirical Risk Minimization Procedure and its Application to Outcome Weighted Learning
di: Giddens, Spencer, et al.
Pubblicazione: (2023)
di: Giddens, Spencer, et al.
Pubblicazione: (2023)
SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights
di: Müller, Lorenz K., et al.
Pubblicazione: (2025)
di: Müller, Lorenz K., et al.
Pubblicazione: (2025)
TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
di: Kim, Jinhee, et al.
Pubblicazione: (2025)
di: Kim, Jinhee, et al.
Pubblicazione: (2025)
A2Q+: Improving Accumulator-Aware Weight Quantization
di: Colbert, Ian, et al.
Pubblicazione: (2024)
di: Colbert, Ian, et al.
Pubblicazione: (2024)
CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs
di: Zhou, Zhaojing, et al.
Pubblicazione: (2025)
di: Zhou, Zhaojing, et al.
Pubblicazione: (2025)
DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick
di: Vali, Mohammad Hassan, et al.
Pubblicazione: (2025)
di: Vali, Mohammad Hassan, et al.
Pubblicazione: (2025)
Leveraging Continuously Differentiable Activation Functions for Learning in Quantized Noisy Environments
di: Shah, Vivswan, et al.
Pubblicazione: (2024)
di: Shah, Vivswan, et al.
Pubblicazione: (2024)
Quantized Convolutional Neural Networks Through the Lens of Partial Differential Equations
di: Ben-Yair, Ido, et al.
Pubblicazione: (2021)
di: Ben-Yair, Ido, et al.
Pubblicazione: (2021)
Hierarchical Vector Quantized Graph Autoencoder with Annealing-Based Code Selection
di: Zeng, Long, et al.
Pubblicazione: (2025)
di: Zeng, Long, et al.
Pubblicazione: (2025)
Differentiable Search for Finding Optimal Quantization Strategy
di: Li, Lianqiang, et al.
Pubblicazione: (2024)
di: Li, Lianqiang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
StatQAT: Statistical Quantizer Optimization for Deep Networks
di: Aktukmak, Mehmet, et al.
Pubblicazione: (2026) -
EfQAT: An Efficient Framework for Quantization-Aware Training
di: Ashkboos, Saleh, et al.
Pubblicazione: (2024) -
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
di: Zhang, Peiyuan, et al.
Pubblicazione: (2026) -
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
di: Gernigon, Cédric, et al.
Pubblicazione: (2024) -
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
di: Ke, Wenjin, et al.
Pubblicazione: (2025)