ArtWeaver: Advanced Dynamic Style Integration via Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Chengming, Hu, Kai, Wang, Qilin, Luo, Donghao, Zhang, Jiangning, Hu, Xiaobin, Fu, Yanwei, Wang, Chengjie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909393066393600
author Xu, Chengming
Hu, Kai
Wang, Qilin
Luo, Donghao
Zhang, Jiangning
Hu, Xiaobin
Fu, Yanwei
Wang, Chengjie
author_facet Xu, Chengming
Hu, Kai
Wang, Qilin
Luo, Donghao
Zhang, Jiangning
Hu, Xiaobin
Fu, Yanwei
Wang, Chengjie
contents Stylized Text-to-Image Generation (STIG) aims to generate images from text prompts and style reference images. In this paper, we present ArtWeaver, a novel framework that leverages pretrained Stable Diffusion (SD) to address challenges such as misinterpreted styles and inconsistent semantics. Our approach introduces two innovative modules: the mixed style descriptor and the dynamic attention adapter. The mixed style descriptor enhances SD by combining content-aware and frequency-disentangled embeddings from CLIP with additional sources that capture global statistics and textual information, thus providing a richer blend of style-related and semantic-related knowledge. To achieve a better balance between adapter capacity and semantic control, the dynamic attention adapter is integrated into the diffusion UNet, dynamically calculating adaptation weights based on the style descriptors. Additionally, we introduce two objective functions to optimize the model alongside the denoising loss, further enhancing semantic and style consistency. Extensive experiments demonstrate the superiority of ArtWeaver over existing methods, producing images with diverse target styles while maintaining the semantic integrity of the text prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15287
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ArtWeaver: Advanced Dynamic Style Integration via Diffusion Model
Xu, Chengming
Hu, Kai
Wang, Qilin
Luo, Donghao
Zhang, Jiangning
Hu, Xiaobin
Fu, Yanwei
Wang, Chengjie
Computer Vision and Pattern Recognition
Stylized Text-to-Image Generation (STIG) aims to generate images from text prompts and style reference images. In this paper, we present ArtWeaver, a novel framework that leverages pretrained Stable Diffusion (SD) to address challenges such as misinterpreted styles and inconsistent semantics. Our approach introduces two innovative modules: the mixed style descriptor and the dynamic attention adapter. The mixed style descriptor enhances SD by combining content-aware and frequency-disentangled embeddings from CLIP with additional sources that capture global statistics and textual information, thus providing a richer blend of style-related and semantic-related knowledge. To achieve a better balance between adapter capacity and semantic control, the dynamic attention adapter is integrated into the diffusion UNet, dynamically calculating adaptation weights based on the style descriptors. Additionally, we introduce two objective functions to optimize the model alongside the denoising loss, further enhancing semantic and style consistency. Extensive experiments demonstrate the superiority of ArtWeaver over existing methods, producing images with diverse target styles while maintaining the semantic integrity of the text prompts.
title ArtWeaver: Advanced Dynamic Style Integration via Diffusion Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.15287