Can Post-Training Quantization Benefit from an Additional QLoRA Integration?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Xiliang, Khasanova, Elena, Chen, Cheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910827785748480
author Zhu, Xiliang
Khasanova, Elena
Chen, Cheng
author_facet Zhu, Xiliang
Khasanova, Elena
Chen, Cheng
contents Large language models (LLMs) have transformed natural language processing but pose significant challenges for real-world deployment. These models necessitate considerable computing resources, which can be costly and frequently unavailable. Model compression techniques such as quantization are often leveraged to alleviate resource demand, but they may have a negative impact on the generation quality. In this study, we explore the integration of 4-bit Post-training Quantization (PTQ) with QLoRA to address these issues. We demonstrate through extensive experiments that this integration outperforms standard PTQ, and in some cases even 16-bit full-parameter fine-tuning on LLMs, validated across proprietary and public datasets with different quantization algorithms. The results demonstrate the efficacy of PTQ-QLoRA integration, offering a viable solution for deploying powerful LLMs in resource-constrained environments without compromising on performance.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10202
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Post-Training Quantization Benefit from an Additional QLoRA Integration?
Zhu, Xiliang
Khasanova, Elena
Chen, Cheng
Computation and Language
Large language models (LLMs) have transformed natural language processing but pose significant challenges for real-world deployment. These models necessitate considerable computing resources, which can be costly and frequently unavailable. Model compression techniques such as quantization are often leveraged to alleviate resource demand, but they may have a negative impact on the generation quality. In this study, we explore the integration of 4-bit Post-training Quantization (PTQ) with QLoRA to address these issues. We demonstrate through extensive experiments that this integration outperforms standard PTQ, and in some cases even 16-bit full-parameter fine-tuning on LLMs, validated across proprietary and public datasets with different quantization algorithms. The results demonstrate the efficacy of PTQ-QLoRA integration, offering a viable solution for deploying powerful LLMs in resource-constrained environments without compromising on performance.
title Can Post-Training Quantization Benefit from an Additional QLoRA Integration?
topic Computation and Language
url https://arxiv.org/abs/2502.10202