IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866912380439494656 |
|---|---|
| author | Guo, Hang Li, Yawei Dai, Tao Xia, Shu-Tao Benini, Luca |
| author_facet | Guo, Hang Li, Yawei Dai, Tao Xia, Shu-Tao Benini, Luca |
| contents | Fine-tuning pre-trained diffusion models under limited budgets has gained great success. In particular, the recent advances that directly fine-tune the quantized weights using Low-rank Adaptation (LoRA) further reduces training costs. Despite these progress, we point out that existing adaptation recipes are not inference-efficient. Specifically, additional post-training quantization (PTQ) on tuned weights is needed during deployment, which results in noticeable performance drop when the bit-width is low. Based on this observation, we introduce IntLoRA, which adapts quantized diffusion models with integer-type low-rank parameters, to include inference efficiency during tuning. Specifically, IntLoRA enables pre-trained weights to remain quantized during training, facilitating fine-tuning on consumer-level GPUs. During inference, IntLoRA weights can be seamlessly merged into pre-trained weights to directly obtain quantized downstream weights without PTQ. Extensive experiments show our IntLoRA achieves significant speedup on both training and inference without losing performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_21759 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion Models Guo, Hang Li, Yawei Dai, Tao Xia, Shu-Tao Benini, Luca Computer Vision and Pattern Recognition Fine-tuning pre-trained diffusion models under limited budgets has gained great success. In particular, the recent advances that directly fine-tune the quantized weights using Low-rank Adaptation (LoRA) further reduces training costs. Despite these progress, we point out that existing adaptation recipes are not inference-efficient. Specifically, additional post-training quantization (PTQ) on tuned weights is needed during deployment, which results in noticeable performance drop when the bit-width is low. Based on this observation, we introduce IntLoRA, which adapts quantized diffusion models with integer-type low-rank parameters, to include inference efficiency during tuning. Specifically, IntLoRA enables pre-trained weights to remain quantized during training, facilitating fine-tuning on consumer-level GPUs. During inference, IntLoRA weights can be seamlessly merged into pre-trained weights to directly obtain quantized downstream weights without PTQ. Extensive experiments show our IntLoRA achieves significant speedup on both training and inference without losing performance. |
| title | IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2410.21759 |