RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Quan, Yau, Chung-Yiu, Wai, Hoi-To, Zhao, Yang Katie, Kang, Dongyeop, Park, Youngsuk, Hong, Mingyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915329856241664
author Wei, Quan
Yau, Chung-Yiu
Wai, Hoi-To
Zhao, Yang Katie
Kang, Dongyeop
Park, Youngsuk
Hong, Mingyi
author_facet Wei, Quan
Yau, Chung-Yiu
Wai, Hoi-To
Zhao, Yang Katie
Kang, Dongyeop
Park, Youngsuk
Hong, Mingyi
contents Supervised fine-tuning is a standard method for adapting pre-trained large language models (LLMs) to downstream tasks. Quantization has been recently studied as a post-training technique for efficient LLM deployment. To obtain quantized fine-tuned LLMs, conventional pipelines would first fine-tune the pre-trained models, followed by post-training quantization. This often yields suboptimal performance as it fails to leverage the synergy between fine-tuning and quantization. To effectively realize low-bit quantization of weights, activations and KV caches in LLMs, we propose an algorithm named Rotated Straight-Through-Estimator (RoSTE), which combines quantization-aware supervised fine-tuning (QA-SFT) with an adaptive rotation strategy that identifies an effective rotation configuration to reduce activation outliers. We provide theoretical insights on RoSTE by analyzing its prediction error when applied to an overparameterized least square quantized training problem. Our findings reveal that the prediction error is directly proportional to the quantization error of the converged weights, which can be effectively managed through an optimized rotation configuration. Experiments on Pythia, Qwen and Llama models of different sizes demonstrate the effectiveness of RoSTE. Compared to existing post-SFT quantization baselines, our method consistently achieves superior performances across various tasks and different LLM architectures. Our code is available at https://github.com/OptimAI-Lab/RoSTE.
format Preprint
id arxiv_https___arxiv_org_abs_2502_09003
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models
Wei, Quan
Yau, Chung-Yiu
Wai, Hoi-To
Zhao, Yang Katie
Kang, Dongyeop
Park, Youngsuk
Hong, Mingyi
Machine Learning
Artificial Intelligence
Supervised fine-tuning is a standard method for adapting pre-trained large language models (LLMs) to downstream tasks. Quantization has been recently studied as a post-training technique for efficient LLM deployment. To obtain quantized fine-tuned LLMs, conventional pipelines would first fine-tune the pre-trained models, followed by post-training quantization. This often yields suboptimal performance as it fails to leverage the synergy between fine-tuning and quantization. To effectively realize low-bit quantization of weights, activations and KV caches in LLMs, we propose an algorithm named Rotated Straight-Through-Estimator (RoSTE), which combines quantization-aware supervised fine-tuning (QA-SFT) with an adaptive rotation strategy that identifies an effective rotation configuration to reduce activation outliers. We provide theoretical insights on RoSTE by analyzing its prediction error when applied to an overparameterized least square quantized training problem. Our findings reveal that the prediction error is directly proportional to the quantization error of the converged weights, which can be effectively managed through an optimized rotation configuration. Experiments on Pythia, Qwen and Llama models of different sizes demonstrate the effectiveness of RoSTE. Compared to existing post-SFT quantization baselines, our method consistently achieves superior performances across various tasks and different LLM architectures. Our code is available at https://github.com/OptimAI-Lab/RoSTE.
title RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.09003