AdaQAT: Adaptive Bit-Width Quantization-Aware Training
Fuente:
arXiv
Guardado en:
| Autores principales: | Gernigon, Cédric, Filip, Silviu-Ioan, Sentieys, Olivier, Coggiola, Clément, Bruno, Mickael |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
por: Zhang, Peiyuan, et al.
Publicado: (2026)
por: Zhang, Peiyuan, et al.
Publicado: (2026)
StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths
por: Chen, Tianyi, et al.
Publicado: (2026)
por: Chen, Tianyi, et al.
Publicado: (2026)
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
por: Chen, Mengzhao, et al.
Publicado: (2024)
por: Chen, Mengzhao, et al.
Publicado: (2024)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
por: Ke, Wenjin, et al.
Publicado: (2025)
por: Ke, Wenjin, et al.
Publicado: (2025)
AdaLoRA-QAT: Adaptive Low-Rank and Quantization-Aware Segmentation
por: Deb, Prantik, et al.
Publicado: (2026)
por: Deb, Prantik, et al.
Publicado: (2026)
Mixed precision accumulation for neural network inference guided by componentwise forward error analysis
por: Arar, El-Mehdi El, et al.
Publicado: (2025)
por: Arar, El-Mehdi El, et al.
Publicado: (2025)
HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs
por: Wang, Guoan, et al.
Publicado: (2026)
por: Wang, Guoan, et al.
Publicado: (2026)
A Stochastic Rounding-Enabled Low-Precision Floating-Point MAC for DNN Training
por: Ali, Sami Ben, et al.
Publicado: (2024)
por: Ali, Sami Ben, et al.
Publicado: (2024)
Ada-Diffuser: Latent-Aware Adaptive Diffusion for Decision-Making
por: Feng, Fan, et al.
Publicado: (2026)
por: Feng, Fan, et al.
Publicado: (2026)
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
por: Lee, Jung Hyun, et al.
Publicado: (2025)
por: Lee, Jung Hyun, et al.
Publicado: (2025)
Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
por: Zhao, Maosen, et al.
Publicado: (2025)
por: Zhao, Maosen, et al.
Publicado: (2025)
Optimizing Large Language Models through Quantization: A Comparative Analysis of PTQ and QAT Techniques
por: Hasan, Jahid
Publicado: (2024)
por: Hasan, Jahid
Publicado: (2024)
MARR: Module-Adaptive Residual Reconstruction for Low-Bit Post-Training Quantization
por: Su, Le, et al.
Publicado: (2026)
por: Su, Le, et al.
Publicado: (2026)
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
por: Lv, Keyu, et al.
Publicado: (2026)
por: Lv, Keyu, et al.
Publicado: (2026)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
por: Zhao, Zhixiong, et al.
Publicado: (2025)
por: Zhao, Zhixiong, et al.
Publicado: (2025)
EfQAT: An Efficient Framework for Quantization-Aware Training
por: Ashkboos, Saleh, et al.
Publicado: (2024)
por: Ashkboos, Saleh, et al.
Publicado: (2024)
Adaptive Width Neural Networks
por: Errica, Federico, et al.
Publicado: (2025)
por: Errica, Federico, et al.
Publicado: (2025)
Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training
por: Varshney, Ayush K., et al.
Publicado: (2026)
por: Varshney, Ayush K., et al.
Publicado: (2026)
Compute-Optimal Quantization-Aware Training
por: Dremov, Aleksandr, et al.
Publicado: (2025)
por: Dremov, Aleksandr, et al.
Publicado: (2025)
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
por: You, Jaeseong, et al.
Publicado: (2024)
por: You, Jaeseong, et al.
Publicado: (2024)
AdaFRUGAL: Adaptive Memory-Efficient Training with Dynamic Control
por: Bui, Quang-Hung, et al.
Publicado: (2025)
por: Bui, Quang-Hung, et al.
Publicado: (2025)
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
por: Nielsen, Jacob, et al.
Publicado: (2025)
por: Nielsen, Jacob, et al.
Publicado: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
por: Su, Zunhai, et al.
Publicado: (2025)
por: Su, Zunhai, et al.
Publicado: (2025)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
por: Zhang, Tianao, et al.
Publicado: (2025)
por: Zhang, Tianao, et al.
Publicado: (2025)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
por: Lee, Deokjae, et al.
Publicado: (2025)
por: Lee, Deokjae, et al.
Publicado: (2025)
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
por: Blumenfeld, Yaniv, et al.
Publicado: (2024)
por: Blumenfeld, Yaniv, et al.
Publicado: (2024)
A Quantized VAE-MLP Botnet Detection Model: A Systematic Evaluation of Quantization-Aware Training and Post-Training Quantization Strategies
por: Wasswa, Hassan, et al.
Publicado: (2025)
por: Wasswa, Hassan, et al.
Publicado: (2025)
BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization
por: Zhao, Jiayu, et al.
Publicado: (2026)
por: Zhao, Jiayu, et al.
Publicado: (2026)
AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size
por: Lu, Guanxi, et al.
Publicado: (2025)
por: Lu, Guanxi, et al.
Publicado: (2025)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
por: Lee, Banseok, et al.
Publicado: (2025)
por: Lee, Banseok, et al.
Publicado: (2025)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
por: Xiao, Guangxuan, et al.
Publicado: (2022)
por: Xiao, Guangxuan, et al.
Publicado: (2022)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
por: Park, Jungwoo, et al.
Publicado: (2025)
por: Park, Jungwoo, et al.
Publicado: (2025)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
por: Wang, Dongwei, et al.
Publicado: (2026)
por: Wang, Dongwei, et al.
Publicado: (2026)
Ada-RS: Adaptive Rejection Sampling for Selective Thinking
por: Ge, Yirou, et al.
Publicado: (2026)
por: Ge, Yirou, et al.
Publicado: (2026)
AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
por: Koh, Woosung, et al.
Publicado: (2025)
por: Koh, Woosung, et al.
Publicado: (2025)
KANtize: Exploring Low-bit Quantization of Kolmogorov-Arnold Networks for Efficient Inference
por: Errabii, Sohaib, et al.
Publicado: (2026)
por: Errabii, Sohaib, et al.
Publicado: (2026)
AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations
por: Tao, Qian, et al.
Publicado: (2024)
por: Tao, Qian, et al.
Publicado: (2024)
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
por: Zhao, Jiaqi, et al.
Publicado: (2025)
por: Zhao, Jiaqi, et al.
Publicado: (2025)
Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
por: Zhang, Xi, et al.
Publicado: (2025)
por: Zhang, Xi, et al.
Publicado: (2025)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
por: Li, Ke, et al.
Publicado: (2026)
por: Li, Ke, et al.
Publicado: (2026)
Ejemplares similares
-
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
por: Zhang, Peiyuan, et al.
Publicado: (2026) -
StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths
por: Chen, Tianyi, et al.
Publicado: (2026) -
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
por: Chen, Mengzhao, et al.
Publicado: (2024) -
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
por: Ke, Wenjin, et al.
Publicado: (2025) -
AdaLoRA-QAT: Adaptive Low-Rank and Quantization-Aware Segmentation
por: Deb, Prantik, et al.
Publicado: (2026)