StyleInject: Parameter Efficient Tuning of Text-to-Image Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Mohan, Bai, Yalong, Yang, Qing, Zhao, Tiejun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911872159055872
author Zhou, Mohan
Bai, Yalong
Yang, Qing
Zhao, Tiejun
author_facet Zhou, Mohan
Bai, Yalong
Yang, Qing
Zhao, Tiejun
contents The ability to fine-tune generative models for text-to-image generation tasks is crucial, particularly facing the complexity involved in accurately interpreting and visualizing textual inputs. While LoRA is efficient for language model adaptation, it often falls short in text-to-image tasks due to the intricate demands of image generation, such as accommodating a broad spectrum of styles and nuances. To bridge this gap, we introduce StyleInject, a specialized fine-tuning approach tailored for text-to-image models. StyleInject comprises multiple parallel low-rank parameter matrices, maintaining the diversity of visual features. It dynamically adapts to varying styles by adjusting the variance of visual features based on the characteristics of the input signal. This approach significantly minimizes the impact on the original model's text-image alignment capabilities while adeptly adapting to various styles in transfer learning. StyleInject proves particularly effective in learning from and enhancing a range of advanced, community-fine-tuned generative models. Our comprehensive experiments, including both small-sample and large-scale data fine-tuning as well as base model distillation, show that StyleInject surpasses traditional LoRA in both text-image semantic consistency and human preference evaluation, all while ensuring greater parameter efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2401_13942
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle StyleInject: Parameter Efficient Tuning of Text-to-Image Diffusion Models
Zhou, Mohan
Bai, Yalong
Yang, Qing
Zhao, Tiejun
Computer Vision and Pattern Recognition
The ability to fine-tune generative models for text-to-image generation tasks is crucial, particularly facing the complexity involved in accurately interpreting and visualizing textual inputs. While LoRA is efficient for language model adaptation, it often falls short in text-to-image tasks due to the intricate demands of image generation, such as accommodating a broad spectrum of styles and nuances. To bridge this gap, we introduce StyleInject, a specialized fine-tuning approach tailored for text-to-image models. StyleInject comprises multiple parallel low-rank parameter matrices, maintaining the diversity of visual features. It dynamically adapts to varying styles by adjusting the variance of visual features based on the characteristics of the input signal. This approach significantly minimizes the impact on the original model's text-image alignment capabilities while adeptly adapting to various styles in transfer learning. StyleInject proves particularly effective in learning from and enhancing a range of advanced, community-fine-tuned generative models. Our comprehensive experiments, including both small-sample and large-scale data fine-tuning as well as base model distillation, show that StyleInject surpasses traditional LoRA in both text-image semantic consistency and human preference evaluation, all while ensuring greater parameter efficiency.
title StyleInject: Parameter Efficient Tuning of Text-to-Image Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.13942