LocRef-Diffusion:Tuning-Free Layout and Appearance-Guided Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deng, Fan, Wu, Yaguang, Yu, Xinyang, Huang, Xiangjun, Yang, Jian, Yan, Guangyu, Xu, Qiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912130982215680
author Deng, Fan
Wu, Yaguang
Yu, Xinyang
Huang, Xiangjun
Yang, Jian
Yan, Guangyu
Xu, Qiang
author_facet Deng, Fan
Wu, Yaguang
Yu, Xinyang
Huang, Xiangjun
Yang, Jian
Yan, Guangyu
Xu, Qiang
contents Recently, text-to-image models based on diffusion have achieved remarkable success in generating high-quality images. However, the challenge of personalized, controllable generation of instances within these images remains an area in need of further development. In this paper, we present LocRef-Diffusion, a novel, tuning-free model capable of personalized customization of multiple instances' appearance and position within an image. To enhance the precision of instance placement, we introduce a Layout-net, which controls instance generation locations by leveraging both explicit instance layout information and an instance region cross-attention module. To improve the appearance fidelity to reference images, we employ an appearance-net that extracts instance appearance features and integrates them into the diffusion model through cross-attention mechanisms. We conducted extensive experiments on the COCO and OpenImages datasets, and the results demonstrate that our proposed method achieves state-of-the-art performance in layout and appearance guided generation.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15252
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LocRef-Diffusion:Tuning-Free Layout and Appearance-Guided Generation
Deng, Fan
Wu, Yaguang
Yu, Xinyang
Huang, Xiangjun
Yang, Jian
Yan, Guangyu
Xu, Qiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Recently, text-to-image models based on diffusion have achieved remarkable success in generating high-quality images. However, the challenge of personalized, controllable generation of instances within these images remains an area in need of further development. In this paper, we present LocRef-Diffusion, a novel, tuning-free model capable of personalized customization of multiple instances' appearance and position within an image. To enhance the precision of instance placement, we introduce a Layout-net, which controls instance generation locations by leveraging both explicit instance layout information and an instance region cross-attention module. To improve the appearance fidelity to reference images, we employ an appearance-net that extracts instance appearance features and integrates them into the diffusion model through cross-attention mechanisms. We conducted extensive experiments on the COCO and OpenImages datasets, and the results demonstrate that our proposed method achieves state-of-the-art performance in layout and appearance guided generation.
title LocRef-Diffusion:Tuning-Free Layout and Appearance-Guided Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.15252