What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lv, Keyu, Zhang, Manyi, Xia, Xiaobo, Ni, Jingchen, Yan, Shannan, Yu, Xianzhi, Hou, Lu, Yuan, Chun, Bai, Haoli
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909996816531456
author Lv, Keyu
Zhang, Manyi
Xia, Xiaobo
Ni, Jingchen
Yan, Shannan
Yu, Xianzhi
Hou, Lu
Yuan, Chun
Bai, Haoli
author_facet Lv, Keyu
Zhang, Manyi
Xia, Xiaobo
Ni, Jingchen
Yan, Shannan
Yu, Xianzhi
Hou, Lu
Yuan, Chun
Bai, Haoli
contents Reasoning models excel at complex tasks such as coding and mathematics, yet their inference is often slow and token-inefficient. To improve the inference efficiency, post-training quantization (PTQ) usually comes with the cost of large accuracy drops, especially for reasoning tasks under low-bit settings. In this study, we present a systematic empirical study of quantization-aware training (QAT) for reasoning models. Our key findings include: (1) Knowledge distillation is a robust objective for reasoning models trained via either supervised fine-tuning or reinforcement learning; (2) PTQ provides a strong initialization for QAT, improving accuracy while reducing training cost; (3) Reinforcement learning remains feasible for quantized models given a viable cold start and yields additional gains; and (4) Aligning the PTQ calibration domain with the QAT training domain accelerates convergence and often improves the final accuracy. Finally, we consolidate these findings into an optimized workflow (Reasoning-QAT), and show that it consistently outperforms state-of-the-art PTQ methods across multiple LLM backbones and reasoning datasets. For instance, on Qwen3-0.6B, it surpasses GPTQ by 44.53% on MATH-500 and consistently recovers performance in the 2-bit regime.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14888
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
Lv, Keyu
Zhang, Manyi
Xia, Xiaobo
Ni, Jingchen
Yan, Shannan
Yu, Xianzhi
Hou, Lu
Yuan, Chun
Bai, Haoli
Machine Learning
Artificial Intelligence
Computation and Language
Reasoning models excel at complex tasks such as coding and mathematics, yet their inference is often slow and token-inefficient. To improve the inference efficiency, post-training quantization (PTQ) usually comes with the cost of large accuracy drops, especially for reasoning tasks under low-bit settings. In this study, we present a systematic empirical study of quantization-aware training (QAT) for reasoning models. Our key findings include: (1) Knowledge distillation is a robust objective for reasoning models trained via either supervised fine-tuning or reinforcement learning; (2) PTQ provides a strong initialization for QAT, improving accuracy while reducing training cost; (3) Reinforcement learning remains feasible for quantized models given a viable cold start and yields additional gains; and (4) Aligning the PTQ calibration domain with the QAT training domain accelerates convergence and often improves the final accuracy. Finally, we consolidate these findings into an optimized workflow (Reasoning-QAT), and show that it consistently outperforms state-of-the-art PTQ methods across multiple LLM backbones and reasoning datasets. For instance, on Qwen3-0.6B, it surpasses GPTQ by 44.53% on MATH-500 and consistently recovers performance in the 2-bit regime.
title What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.14888