StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Tianyi, Chen, Sihan, Qu, Xiaoyi, Zhao, Dan, Yan, Ruomei, Ko, Jongwoo, Liang, Luming, Cameron, Pashmina |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
von: Chen, Sihan, et al.
Veröffentlicht: (2025)
von: Chen, Sihan, et al.
Veröffentlicht: (2025)
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
von: Ko, Jongwoo, et al.
Veröffentlicht: (2026)
von: Ko, Jongwoo, et al.
Veröffentlicht: (2026)
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
von: Gernigon, Cédric, et al.
Veröffentlicht: (2024)
von: Gernigon, Cédric, et al.
Veröffentlicht: (2024)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2026)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
von: Ko, Jongwoo, et al.
Veröffentlicht: (2025)
von: Ko, Jongwoo, et al.
Veröffentlicht: (2025)
AdaLoRA-QAT: Adaptive Low-Rank and Quantization-Aware Segmentation
von: Deb, Prantik, et al.
Veröffentlicht: (2026)
von: Deb, Prantik, et al.
Veröffentlicht: (2026)
QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models
von: Liu, Jing, et al.
Veröffentlicht: (2023)
von: Liu, Jing, et al.
Veröffentlicht: (2023)
HESSO: Towards Automatic Efficient and User Friendly Any Neural Network Training and Pruning
von: Chen, Tianyi, et al.
Veröffentlicht: (2024)
von: Chen, Tianyi, et al.
Veröffentlicht: (2024)
PQS (Prune, Quantize, and Sort): Low-Bitwidth Accumulation of Dot Products in Neural Network Computations
von: Natesh, Vikas, et al.
Veröffentlicht: (2025)
von: Natesh, Vikas, et al.
Veröffentlicht: (2025)
FraQAT: Quantization Aware Training with Fractional bits
von: Morreale, Luca, et al.
Veröffentlicht: (2025)
von: Morreale, Luca, et al.
Veröffentlicht: (2025)
EfQAT: An Efficient Framework for Quantization-Aware Training
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)
DistiLLM: Towards Streamlined Distillation for Large Language Models
von: Ko, Jongwoo, et al.
Veröffentlicht: (2024)
von: Ko, Jongwoo, et al.
Veröffentlicht: (2024)
Automatic Joint Structured Pruning and Quantization for Efficient Neural Network Training and Compression
von: Qu, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Qu, Xiaoyi, et al.
Veröffentlicht: (2025)
CUA-Skill: Develop Skills for Computer Using Agent
von: Chen, Tianyi, et al.
Veröffentlicht: (2026)
von: Chen, Tianyi, et al.
Veröffentlicht: (2026)
Bitwidth-Specific Logarithmic Arithmetic for Future Hardware-Accelerated Training
von: Hamad, Hassan, et al.
Veröffentlicht: (2025)
von: Hamad, Hassan, et al.
Veröffentlicht: (2025)
MF-QAT: Multi-Format Quantization-Aware Training for Elastic Inference
von: Xu, Zifei, et al.
Veröffentlicht: (2026)
von: Xu, Zifei, et al.
Veröffentlicht: (2026)
Low-Bitwidth Floating Point Quantization for Efficient High-Quality Diffusion Models
von: Chen, Cheng, et al.
Veröffentlicht: (2024)
von: Chen, Cheng, et al.
Veröffentlicht: (2024)
QaRL: Rollout-Aligned Quantization-Aware RL for Fast and Stable Training under Training--Inference Mismatch
von: Gu, Hao, et al.
Veröffentlicht: (2026)
von: Gu, Hao, et al.
Veröffentlicht: (2026)
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2025)
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2025)
Low-Rank Quantization-Aware Training for LLMs
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training
von: Emadi, Seyed Morteza
Veröffentlicht: (2026)
von: Emadi, Seyed Morteza
Veröffentlicht: (2026)
Power-of-Two Quantization-Aware-Training (PoT-QAT) in Large Language Models (LLMs)
von: Elgenedy, Mahmoud
Veröffentlicht: (2026)
von: Elgenedy, Mahmoud
Veröffentlicht: (2026)
CR-QAT: Curriculum Relational Quantization-Aware Training for Open-Vocabulary Object Detection
von: Park, Jinyeong, et al.
Veröffentlicht: (2026)
von: Park, Jinyeong, et al.
Veröffentlicht: (2026)
Optimizing Large Language Models through Quantization: A Comparative Analysis of PTQ and QAT Techniques
von: Hasan, Jahid
Veröffentlicht: (2024)
von: Hasan, Jahid
Veröffentlicht: (2024)
Bayesian Principles Improve Prompt Learning In Vision-Language Models
von: Kim, Mingyu, et al.
Veröffentlicht: (2025)
von: Kim, Mingyu, et al.
Veröffentlicht: (2025)
StableQuant: Layer Adaptive Post-Training Quantization for Speech Foundation Models
von: Hong, Yeona, et al.
Veröffentlicht: (2025)
von: Hong, Yeona, et al.
Veröffentlicht: (2025)
Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
von: Zhao, Maosen, et al.
Veröffentlicht: (2025)
von: Zhao, Maosen, et al.
Veröffentlicht: (2025)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
von: Luo, Yingsong, et al.
Veröffentlicht: (2024)
von: Luo, Yingsong, et al.
Veröffentlicht: (2024)
Progressive Fine-to-Coarse Reconstruction for Accurate Low-Bit Post-Training Quantization in Vision Transformers
von: Ding, Rui, et al.
Veröffentlicht: (2024)
von: Ding, Rui, et al.
Veröffentlicht: (2024)
Towards Difficulty-Agnostic Efficient Transfer Learning for Vision-Language Models
von: Yang, Yongjin, et al.
Veröffentlicht: (2023)
von: Yang, Yongjin, et al.
Veröffentlicht: (2023)
Unexplainability of Artificial Intelligence Judgments in Kant's Perspective
von: Seo, Jongwoo
Veröffentlicht: (2024)
von: Seo, Jongwoo
Veröffentlicht: (2024)
Cat-AIR: Content and Task-Aware All-in-One Image Restoration
von: Jiang, Jiachen, et al.
Veröffentlicht: (2025)
von: Jiang, Jiachen, et al.
Veröffentlicht: (2025)
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
Aspect-Aware MOOC Recommendation in a Heterogeneous Network
von: Chu, Seongyeub, et al.
Veröffentlicht: (2026)
von: Chu, Seongyeub, et al.
Veröffentlicht: (2026)
Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
von: Xu, Binxing, et al.
Veröffentlicht: (2026)
von: Xu, Binxing, et al.
Veröffentlicht: (2026)
SASQ: Static Activation Scaling for Quantization-Aware Training in Large Language Models
von: Mao, Shizhuo, et al.
Veröffentlicht: (2025)
von: Mao, Shizhuo, et al.
Veröffentlicht: (2025)
Compute-Optimal Quantization-Aware Training
von: Dremov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Dremov, Aleksandr, et al.
Veröffentlicht: (2025)
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
von: You, Jaeseong, et al.
Veröffentlicht: (2024)
von: You, Jaeseong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
von: Chen, Sihan, et al.
Veröffentlicht: (2025) -
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
von: Ko, Jongwoo, et al.
Veröffentlicht: (2026) -
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024) -
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
von: Gernigon, Cédric, et al.
Veröffentlicht: (2024) -
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2026)