End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tan, Qitao, Song, Xiaoying, Lu, Jin, Li, Guoming, Liu, Jun, Hong, Lingzi, Ding, Caiwen, Li, Jundong, Zhai, Xiaoming, Huang, Shaoyi, Niu, Wei, Yuan, Geng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915520967606272
author Tan, Qitao
Song, Xiaoying
Lu, Jin
Li, Guoming
Liu, Jun
Hong, Lingzi
Ding, Caiwen
Li, Jundong
Zhai, Xiaoming
Huang, Shaoyi
Niu, Wei
Yuan, Geng
author_facet Tan, Qitao
Song, Xiaoying
Lu, Jin
Li, Guoming
Liu, Jun
Hong, Lingzi
Ding, Caiwen
Li, Jundong
Zhai, Xiaoming
Huang, Shaoyi
Niu, Wei
Yuan, Geng
contents Quantization is an effective technique to reduce the deployment cost of large language models (LLMs), and post-training quantization (PTQ) has been widely studied due to its efficiency. However, existing PTQ methods are limited by their inability to fine-tune model parameters and often suffer significant accuracy loss in low-bit scenarios. Quantization-aware training (QAT) provides a more principled solution, but its reliance on backpropagation incurs prohibitive memory costs, limiting its practicality for LLM deployment. To address these challenges, we propose ZeroQAT, a zeroth-order optimization-based QAT framework that supports both weight and activation quantization. ZeroQAT leverages forward-only gradient estimation to eliminate backpropagation, substantially reducing computational and memory overhead while retaining the benefits of end-to-end optimization. We further introduce a lightweight variant of ZeroQAT for quantized fine-tuning, which freezes and pre-quantizes most parameters to further cut memory usage. Experiments show that ZeroQAT consistently outperforms representative PTQ and QAT baselines while requiring significantly less memory. For example, ZeroQAT enables fine-tuning of a 13B model at extremely low bit-widths (e.g., 2-4 bits) on a single 8GB GPU, and even allows fine-tuning a 6.7B model on a OnePlus 12 smartphone, demonstrating its practicality for end-to-end QAT on resource-limited edge devices.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00031
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
Tan, Qitao
Song, Xiaoying
Lu, Jin
Li, Guoming
Liu, Jun
Hong, Lingzi
Ding, Caiwen
Li, Jundong
Zhai, Xiaoming
Huang, Shaoyi
Niu, Wei
Yuan, Geng
Machine Learning
Artificial Intelligence
Quantization is an effective technique to reduce the deployment cost of large language models (LLMs), and post-training quantization (PTQ) has been widely studied due to its efficiency. However, existing PTQ methods are limited by their inability to fine-tune model parameters and often suffer significant accuracy loss in low-bit scenarios. Quantization-aware training (QAT) provides a more principled solution, but its reliance on backpropagation incurs prohibitive memory costs, limiting its practicality for LLM deployment. To address these challenges, we propose ZeroQAT, a zeroth-order optimization-based QAT framework that supports both weight and activation quantization. ZeroQAT leverages forward-only gradient estimation to eliminate backpropagation, substantially reducing computational and memory overhead while retaining the benefits of end-to-end optimization. We further introduce a lightweight variant of ZeroQAT for quantized fine-tuning, which freezes and pre-quantizes most parameters to further cut memory usage. Experiments show that ZeroQAT consistently outperforms representative PTQ and QAT baselines while requiring significantly less memory. For example, ZeroQAT enables fine-tuning of a 13B model at extremely low bit-widths (e.g., 2-4 bits) on a single 8GB GPU, and even allows fine-tuning a 6.7B model on a OnePlus 12 smartphone, demonstrating its practicality for end-to-end QAT on resource-limited edge devices.
title End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.00031