StyleAdapter: A Unified Stylized Image Generation Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhouxia, Wang, Xintao, Xie, Liangbin, Qi, Zhongang, Shan, Ying, Wang, Wenping, Luo, Ping
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913567438012416
author Wang, Zhouxia
Wang, Xintao
Xie, Liangbin
Qi, Zhongang
Shan, Ying
Wang, Wenping
Luo, Ping
author_facet Wang, Zhouxia
Wang, Xintao
Xie, Liangbin
Qi, Zhongang
Shan, Ying
Wang, Wenping
Luo, Ping
contents This work focuses on generating high-quality images with specific style of reference images and content of provided textual descriptions. Current leading algorithms, i.e., DreamBooth and LoRA, require fine-tuning for each style, leading to time-consuming and computationally expensive processes. In this work, we propose StyleAdapter, a unified stylized image generation model capable of producing a variety of stylized images that match both the content of a given prompt and the style of reference images, without the need for per-style fine-tuning. It introduces a two-path cross-attention (TPCA) module to separately process style information and textual prompt, which cooperate with a semantic suppressing vision model (SSVM) to suppress the semantic content of style images. In this way, it can ensure that the prompt maintains control over the content of the generated images, while also mitigating the negative impact of semantic information in style references. This results in the content of the generated image adhering to the prompt, and its style aligning with the style references. Besides, our StyleAdapter can be integrated with existing controllable synthesis methods, such as T2I-adapter and ControlNet, to attain a more controllable and stable generation process. Extensive experiments demonstrate the superiority of our method over previous works.
format Preprint
id arxiv_https___arxiv_org_abs_2309_01770
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle StyleAdapter: A Unified Stylized Image Generation Model
Wang, Zhouxia
Wang, Xintao
Xie, Liangbin
Qi, Zhongang
Shan, Ying
Wang, Wenping
Luo, Ping
Computer Vision and Pattern Recognition
This work focuses on generating high-quality images with specific style of reference images and content of provided textual descriptions. Current leading algorithms, i.e., DreamBooth and LoRA, require fine-tuning for each style, leading to time-consuming and computationally expensive processes. In this work, we propose StyleAdapter, a unified stylized image generation model capable of producing a variety of stylized images that match both the content of a given prompt and the style of reference images, without the need for per-style fine-tuning. It introduces a two-path cross-attention (TPCA) module to separately process style information and textual prompt, which cooperate with a semantic suppressing vision model (SSVM) to suppress the semantic content of style images. In this way, it can ensure that the prompt maintains control over the content of the generated images, while also mitigating the negative impact of semantic information in style references. This results in the content of the generated image adhering to the prompt, and its style aligning with the style references. Besides, our StyleAdapter can be integrated with existing controllable synthesis methods, such as T2I-adapter and ControlNet, to attain a more controllable and stable generation process. Extensive experiments demonstrate the superiority of our method over previous works.
title StyleAdapter: A Unified Stylized Image Generation Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2309.01770