CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deng, Yanxia, Zhang, Aozhong, Gurses, Selcuk, Wang, Naigang, Yang, Zi, Yin, Penghang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916897168031744
author Deng, Yanxia
Zhang, Aozhong
Gurses, Selcuk
Wang, Naigang
Yang, Zi
Yin, Penghang
author_facet Deng, Yanxia
Zhang, Aozhong
Gurses, Selcuk
Wang, Naigang
Yang, Zi
Yin, Penghang
contents Fine-tuning large language models (LLMs) using low-rank adaptation (LoRA) has become a highly efficient approach for downstream tasks, particularly in scenarios with limited computational resources. However, applying LoRA techniques to quantized LLMs poses unique challenges due to the reduced representational precision of quantized weights. In this paper, we introduce CLoQ (Calibrated LoRA initialization for Quantized LLMs), a simplistic initialization strategy designed to overcome these challenges. Our approach focuses on minimizing the layer-wise discrepancy between the original LLM and its quantized counterpart with LoRA components during initialization. By leveraging a small calibration dataset, CLoQ quantizes a pre-trained LLM and determines the optimal LoRA components for each layer, ensuring a strong foundation for subsequent fine-tuning. A key contribution of this work is a novel theoretical result that enables the accurate and closed-form construction of these optimal LoRA components. We validate the efficacy of CLoQ across multiple tasks such as language generation, arithmetic reasoning, and commonsense reasoning, demonstrating that it consistently outperforms existing LoRA fine-tuning methods for quantized LLMs, especially at ultra low-bit widths.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18475
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization
Deng, Yanxia
Zhang, Aozhong
Gurses, Selcuk
Wang, Naigang
Yang, Zi
Yin, Penghang
Machine Learning
Artificial Intelligence
Fine-tuning large language models (LLMs) using low-rank adaptation (LoRA) has become a highly efficient approach for downstream tasks, particularly in scenarios with limited computational resources. However, applying LoRA techniques to quantized LLMs poses unique challenges due to the reduced representational precision of quantized weights. In this paper, we introduce CLoQ (Calibrated LoRA initialization for Quantized LLMs), a simplistic initialization strategy designed to overcome these challenges. Our approach focuses on minimizing the layer-wise discrepancy between the original LLM and its quantized counterpart with LoRA components during initialization. By leveraging a small calibration dataset, CLoQ quantizes a pre-trained LLM and determines the optimal LoRA components for each layer, ensuring a strong foundation for subsequent fine-tuning. A key contribution of this work is a novel theoretical result that enables the accurate and closed-form construction of these optimal LoRA components. We validate the efficacy of CLoQ across multiple tasks such as language generation, arithmetic reasoning, and commonsense reasoning, demonstrating that it consistently outperforms existing LoRA fine-tuning methods for quantized LLMs, especially at ultra low-bit widths.
title CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2501.18475