Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Jung Hyun, Shin, Seungjae, Kim, Vinnam, You, Jaeseong, Chen, An |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing Large Language Models through Quantization: A Comparative Analysis of PTQ and QAT Techniques
by: Hasan, Jahid
Published: (2024)
by: Hasan, Jahid
Published: (2024)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
by: Lee, Jung Hyun, et al.
Published: (2023)
by: Lee, Jung Hyun, et al.
Published: (2023)
Pack-PTQ: Advancing Post-training Quantization of Neural Networks by Pack-wise Reconstruction
by: Li, Changjun, et al.
Published: (2025)
by: Li, Changjun, et al.
Published: (2025)
FraQAT: Quantization Aware Training with Fractional bits
by: Morreale, Luca, et al.
Published: (2025)
by: Morreale, Luca, et al.
Published: (2025)
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
by: You, Jaeseong, et al.
Published: (2024)
by: You, Jaeseong, et al.
Published: (2024)
Two-Stage Grid Optimization for Group-wise Quantization of LLMs
by: Kim, Junhan, et al.
Published: (2026)
by: Kim, Junhan, et al.
Published: (2026)
StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths
by: Chen, Tianyi, et al.
Published: (2026)
by: Chen, Tianyi, et al.
Published: (2026)
Realizing Unaligned Block-wise Pruning for DNN Acceleration on Mobile Devices
by: Lee, Hayun, et al.
Published: (2024)
by: Lee, Hayun, et al.
Published: (2024)
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
by: Gernigon, Cédric, et al.
Published: (2024)
by: Gernigon, Cédric, et al.
Published: (2024)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
by: Zhang, Peiyuan, et al.
Published: (2026)
by: Zhang, Peiyuan, et al.
Published: (2026)
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
by: Chen, Mengzhao, et al.
Published: (2024)
by: Chen, Mengzhao, et al.
Published: (2024)
LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs
by: Lee, Jung Hyun, et al.
Published: (2026)
by: Lee, Jung Hyun, et al.
Published: (2026)
Safety-Preserving PTQ via Contrastive Alignment Loss
by: Wee, Sunghyun, et al.
Published: (2025)
by: Wee, Sunghyun, et al.
Published: (2025)
PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation Models
by: Liu, Xuewen, et al.
Published: (2026)
by: Liu, Xuewen, et al.
Published: (2026)
GranQ: Efficient Channel-wise Quantization via Vectorized Pre-Scaling for Zero-Shot QAT
by: Hong, Inpyo, et al.
Published: (2025)
by: Hong, Inpyo, et al.
Published: (2025)
EuraGovExam: A Multilingual Multimodal Benchmark from Real-World Civil Service Exams
by: Kim, Jaeseong, et al.
Published: (2026)
by: Kim, Jaeseong, et al.
Published: (2026)
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
by: Xia, Haojun, et al.
Published: (2025)
by: Xia, Haojun, et al.
Published: (2025)
DeltaDQ: Ultra-High Delta Compression for Fine-Tuned LLMs via Group-wise Dropout and Separate Quantization
by: Jiang, Yanfeng, et al.
Published: (2024)
by: Jiang, Yanfeng, et al.
Published: (2024)
PTQ4VM: Post-Training Quantization for Visual Mamba
by: Cho, Younghyun, et al.
Published: (2024)
by: Cho, Younghyun, et al.
Published: (2024)
AdaLoRA-QAT: Adaptive Low-Rank and Quantization-Aware Segmentation
by: Deb, Prantik, et al.
Published: (2026)
by: Deb, Prantik, et al.
Published: (2026)
Accurate Block Quantization in LLMs with Outliers
by: Trukhanov, Nikita, et al.
Published: (2024)
by: Trukhanov, Nikita, et al.
Published: (2024)
Block-wise Codeword Embedding for Reliable Multi-bit Text Watermarking
by: Kim, Joeun, et al.
Published: (2026)
by: Kim, Joeun, et al.
Published: (2026)
BATQuant: Outlier-resilient MXFP4 Quantization via Learnable Block-wise Optimization
by: Li, Ji-Fu, et al.
Published: (2026)
by: Li, Ji-Fu, et al.
Published: (2026)
A Case Study of Selected PTQ Baselines for Reasoning LLMs on Ascend NPU
by: Luo, Yuchen, et al.
Published: (2026)
by: Luo, Yuchen, et al.
Published: (2026)
Rainbow Padding: Mitigating Early Termination in Instruction-Tuned Diffusion LLMs
by: Kim, Bumjun, et al.
Published: (2025)
by: Kim, Bumjun, et al.
Published: (2025)
Edge-FIT: Federated Instruction Tuning of Quantized LLMs for Privacy-Preserving Smart Home Environments
by: Venkatesh, Vinay, et al.
Published: (2025)
by: Venkatesh, Vinay, et al.
Published: (2025)
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
by: Lee, Geonho, et al.
Published: (2024)
by: Lee, Geonho, et al.
Published: (2024)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
by: Ke, Wenjin, et al.
Published: (2025)
by: Ke, Wenjin, et al.
Published: (2025)
Channel-wise Vector Quantization
by: Song, Wei, et al.
Published: (2026)
by: Song, Wei, et al.
Published: (2026)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
by: Yan, Xianglong, et al.
Published: (2026)
by: Yan, Xianglong, et al.
Published: (2026)
SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
by: Xia, Junhao, et al.
Published: (2025)
by: Xia, Junhao, et al.
Published: (2025)
Semantic Guidance Tuning for Text-To-Image Diffusion Models
by: Kang, Hyun, et al.
Published: (2023)
by: Kang, Hyun, et al.
Published: (2023)
Efficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise Replacement
by: Yang, Shu, et al.
Published: (2025)
by: Yang, Shu, et al.
Published: (2025)
InstructPipe: Generating Visual Blocks Pipelines with Human Instructions and LLMs
by: Zhou, Zhongyi, et al.
Published: (2023)
by: Zhou, Zhongyi, et al.
Published: (2023)
4bit-Quantization in Vector-Embedding for RAG
by: Jeong, Taehee
Published: (2025)
by: Jeong, Taehee
Published: (2025)
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
by: Zhao, Jiaqi, et al.
Published: (2025)
by: Zhao, Jiaqi, et al.
Published: (2025)
Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding
by: Yun, Taewon, et al.
Published: (2026)
by: Yun, Taewon, et al.
Published: (2026)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
UMIE: Unified Multimodal Information Extraction with Instruction Tuning
by: Sun, Lin, et al.
Published: (2024)
by: Sun, Lin, et al.
Published: (2024)
TiVaT: A Transformer with a Single Unified Mechanism for Capturing Asynchronous Dependencies in Multivariate Time Series Forecasting
by: Ha, Junwoo, et al.
Published: (2024)
by: Ha, Junwoo, et al.
Published: (2024)
Similar Items
-
Optimizing Large Language Models through Quantization: A Comparative Analysis of PTQ and QAT Techniques
by: Hasan, Jahid
Published: (2024) -
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
by: Lee, Jung Hyun, et al.
Published: (2023) -
Pack-PTQ: Advancing Post-training Quantization of Neural Networks by Pack-wise Reconstruction
by: Li, Changjun, et al.
Published: (2025) -
FraQAT: Quantization Aware Training with Fractional bits
by: Morreale, Luca, et al.
Published: (2025) -
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
by: You, Jaeseong, et al.
Published: (2024)