CDM-QTA: Quantized Training Acceleration for Efficient LoRA Fine-Tuning of Diffusion Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lu, Jinming, She, Minghao, Mao, Wendong, Wang, Zhongfeng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916685172178944
author Lu, Jinming
She, Minghao
Mao, Wendong
Wang, Zhongfeng
author_facet Lu, Jinming
She, Minghao
Mao, Wendong
Wang, Zhongfeng
contents Fine-tuning large diffusion models for custom applications demands substantial power and time, which poses significant challenges for efficient implementation on mobile devices. In this paper, we develop a novel training accelerator specifically for Low-Rank Adaptation (LoRA) of diffusion models, aiming to streamline the process and reduce computational complexity. By leveraging a fully quantized training scheme for LoRA fine-tuning, we achieve substantial reductions in memory usage and power consumption while maintaining high model fidelity. The proposed accelerator features flexible dataflow, enabling high utilization for irregular and variable tensor shapes during the LoRA process. Experimental results show up to 1.81x training speedup and 5.50x energy efficiency improvements compared to the baseline, with minimal impact on image generation quality.
format Preprint
id arxiv_https___arxiv_org_abs_2504_07998
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CDM-QTA: Quantized Training Acceleration for Efficient LoRA Fine-Tuning of Diffusion Model
Lu, Jinming
She, Minghao
Mao, Wendong
Wang, Zhongfeng
Graphics
Artificial Intelligence
Hardware Architecture
Computer Vision and Pattern Recognition
Fine-tuning large diffusion models for custom applications demands substantial power and time, which poses significant challenges for efficient implementation on mobile devices. In this paper, we develop a novel training accelerator specifically for Low-Rank Adaptation (LoRA) of diffusion models, aiming to streamline the process and reduce computational complexity. By leveraging a fully quantized training scheme for LoRA fine-tuning, we achieve substantial reductions in memory usage and power consumption while maintaining high model fidelity. The proposed accelerator features flexible dataflow, enabling high utilization for irregular and variable tensor shapes during the LoRA process. Experimental results show up to 1.81x training speedup and 5.50x energy efficiency improvements compared to the baseline, with minimal impact on image generation quality.
title CDM-QTA: Quantized Training Acceleration for Efficient LoRA Fine-Tuning of Diffusion Model
topic Graphics
Artificial Intelligence
Hardware Architecture
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.07998