End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915520967606272 |
|---|---|
| author | Tan, Qitao Song, Xiaoying Lu, Jin Li, Guoming Liu, Jun Hong, Lingzi Ding, Caiwen Li, Jundong Zhai, Xiaoming Huang, Shaoyi Niu, Wei Yuan, Geng |
| author_facet | Tan, Qitao Song, Xiaoying Lu, Jin Li, Guoming Liu, Jun Hong, Lingzi Ding, Caiwen Li, Jundong Zhai, Xiaoming Huang, Shaoyi Niu, Wei Yuan, Geng |
| contents | Quantization is an effective technique to reduce the deployment cost of large language models (LLMs), and post-training quantization (PTQ) has been widely studied due to its efficiency. However, existing PTQ methods are limited by their inability to fine-tune model parameters and often suffer significant accuracy loss in low-bit scenarios. Quantization-aware training (QAT) provides a more principled solution, but its reliance on backpropagation incurs prohibitive memory costs, limiting its practicality for LLM deployment. To address these challenges, we propose ZeroQAT, a zeroth-order optimization-based QAT framework that supports both weight and activation quantization. ZeroQAT leverages forward-only gradient estimation to eliminate backpropagation, substantially reducing computational and memory overhead while retaining the benefits of end-to-end optimization. We further introduce a lightweight variant of ZeroQAT for quantized fine-tuning, which freezes and pre-quantizes most parameters to further cut memory usage. Experiments show that ZeroQAT consistently outperforms representative PTQ and QAT baselines while requiring significantly less memory. For example, ZeroQAT enables fine-tuning of a 13B model at extremely low bit-widths (e.g., 2-4 bits) on a single 8GB GPU, and even allows fine-tuning a 6.7B model on a OnePlus 12 smartphone, demonstrating its practicality for end-to-end QAT on resource-limited edge devices. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_00031 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost Tan, Qitao Song, Xiaoying Lu, Jin Li, Guoming Liu, Jun Hong, Lingzi Ding, Caiwen Li, Jundong Zhai, Xiaoming Huang, Shaoyi Niu, Wei Yuan, Geng Machine Learning Artificial Intelligence Quantization is an effective technique to reduce the deployment cost of large language models (LLMs), and post-training quantization (PTQ) has been widely studied due to its efficiency. However, existing PTQ methods are limited by their inability to fine-tune model parameters and often suffer significant accuracy loss in low-bit scenarios. Quantization-aware training (QAT) provides a more principled solution, but its reliance on backpropagation incurs prohibitive memory costs, limiting its practicality for LLM deployment. To address these challenges, we propose ZeroQAT, a zeroth-order optimization-based QAT framework that supports both weight and activation quantization. ZeroQAT leverages forward-only gradient estimation to eliminate backpropagation, substantially reducing computational and memory overhead while retaining the benefits of end-to-end optimization. We further introduce a lightweight variant of ZeroQAT for quantized fine-tuning, which freezes and pre-quantizes most parameters to further cut memory usage. Experiments show that ZeroQAT consistently outperforms representative PTQ and QAT baselines while requiring significantly less memory. For example, ZeroQAT enables fine-tuning of a 13B model at extremely low bit-widths (e.g., 2-4 bits) on a single 8GB GPU, and even allows fine-tuning a 6.7B model on a OnePlus 12 smartphone, demonstrating its practicality for end-to-end QAT on resource-limited edge devices. |
| title | End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2509.00031 |