IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Guo, Hang, Li, Yawei, Dai, Tao, Xia, Shu-Tao, Benini, Luca
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912380439494656
author Guo, Hang
Li, Yawei
Dai, Tao
Xia, Shu-Tao
Benini, Luca
author_facet Guo, Hang
Li, Yawei
Dai, Tao
Xia, Shu-Tao
Benini, Luca
contents Fine-tuning pre-trained diffusion models under limited budgets has gained great success. In particular, the recent advances that directly fine-tune the quantized weights using Low-rank Adaptation (LoRA) further reduces training costs. Despite these progress, we point out that existing adaptation recipes are not inference-efficient. Specifically, additional post-training quantization (PTQ) on tuned weights is needed during deployment, which results in noticeable performance drop when the bit-width is low. Based on this observation, we introduce IntLoRA, which adapts quantized diffusion models with integer-type low-rank parameters, to include inference efficiency during tuning. Specifically, IntLoRA enables pre-trained weights to remain quantized during training, facilitating fine-tuning on consumer-level GPUs. During inference, IntLoRA weights can be seamlessly merged into pre-trained weights to directly obtain quantized downstream weights without PTQ. Extensive experiments show our IntLoRA achieves significant speedup on both training and inference without losing performance.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21759
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion Models
Guo, Hang
Li, Yawei
Dai, Tao
Xia, Shu-Tao
Benini, Luca
Computer Vision and Pattern Recognition
Fine-tuning pre-trained diffusion models under limited budgets has gained great success. In particular, the recent advances that directly fine-tune the quantized weights using Low-rank Adaptation (LoRA) further reduces training costs. Despite these progress, we point out that existing adaptation recipes are not inference-efficient. Specifically, additional post-training quantization (PTQ) on tuned weights is needed during deployment, which results in noticeable performance drop when the bit-width is low. Based on this observation, we introduce IntLoRA, which adapts quantized diffusion models with integer-type low-rank parameters, to include inference efficiency during tuning. Specifically, IntLoRA enables pre-trained weights to remain quantized during training, facilitating fine-tuning on consumer-level GPUs. During inference, IntLoRA weights can be seamlessly merged into pre-trained weights to directly obtain quantized downstream weights without PTQ. Extensive experiments show our IntLoRA achieves significant speedup on both training and inference without losing performance.
title IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.21759