VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Shaojin, Ding, Fei, Huang, Mengqi, Liu, Wei, He, Qian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913629244227584
author Wu, Shaojin
Ding, Fei
Huang, Mengqi
Liu, Wei
He, Qian
author_facet Wu, Shaojin
Ding, Fei
Huang, Mengqi
Liu, Wei
He, Qian
contents While diffusion models show extraordinary talents in text-to-image generation, they may still fail to generate highly aesthetic images. More specifically, there is still a gap between the generated images and the real-world aesthetic images in finer-grained dimensions including color, lighting, composition, etc. In this paper, we propose Cross-Attention Value Mixing Control (VMix) Adapter, a plug-and-play aesthetics adapter, to upgrade the quality of generated images while maintaining generality across visual concepts by (1) disentangling the input text prompt into the content description and aesthetic description by the initialization of aesthetic embedding, and (2) integrating aesthetic conditions into the denoising process through value-mixed cross-attention, with the network connected by zero-initialized linear layers. Our key insight is to enhance the aesthetic presentation of existing diffusion models by designing a superior condition control method, all while preserving the image-text alignment. Through our meticulous design, VMix is flexible enough to be applied to community models for better visual performance without retraining. To validate the effectiveness of our method, we conducted extensive experiments, showing that VMix outperforms other state-of-the-art methods and is compatible with other community modules (e.g., LoRA, ControlNet, and IPAdapter) for image generation. The project page is https://vmix-diffusion.github.io/VMix/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20800
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control
Wu, Shaojin
Ding, Fei
Huang, Mengqi
Liu, Wei
He, Qian
Computer Vision and Pattern Recognition
While diffusion models show extraordinary talents in text-to-image generation, they may still fail to generate highly aesthetic images. More specifically, there is still a gap between the generated images and the real-world aesthetic images in finer-grained dimensions including color, lighting, composition, etc. In this paper, we propose Cross-Attention Value Mixing Control (VMix) Adapter, a plug-and-play aesthetics adapter, to upgrade the quality of generated images while maintaining generality across visual concepts by (1) disentangling the input text prompt into the content description and aesthetic description by the initialization of aesthetic embedding, and (2) integrating aesthetic conditions into the denoising process through value-mixed cross-attention, with the network connected by zero-initialized linear layers. Our key insight is to enhance the aesthetic presentation of existing diffusion models by designing a superior condition control method, all while preserving the image-text alignment. Through our meticulous design, VMix is flexible enough to be applied to community models for better visual performance without retraining. To validate the effectiveness of our method, we conducted extensive experiments, showing that VMix outperforms other state-of-the-art methods and is compatible with other community modules (e.g., LoRA, ControlNet, and IPAdapter) for image generation. The project page is https://vmix-diffusion.github.io/VMix/.
title VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.20800