LocRef-Diffusion:Tuning-Free Layout and Appearance-Guided Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912130982215680 |
|---|---|
| author | Deng, Fan Wu, Yaguang Yu, Xinyang Huang, Xiangjun Yang, Jian Yan, Guangyu Xu, Qiang |
| author_facet | Deng, Fan Wu, Yaguang Yu, Xinyang Huang, Xiangjun Yang, Jian Yan, Guangyu Xu, Qiang |
| contents | Recently, text-to-image models based on diffusion have achieved remarkable success in generating high-quality images. However, the challenge of personalized, controllable generation of instances within these images remains an area in need of further development. In this paper, we present LocRef-Diffusion, a novel, tuning-free model capable of personalized customization of multiple instances' appearance and position within an image. To enhance the precision of instance placement, we introduce a Layout-net, which controls instance generation locations by leveraging both explicit instance layout information and an instance region cross-attention module. To improve the appearance fidelity to reference images, we employ an appearance-net that extracts instance appearance features and integrates them into the diffusion model through cross-attention mechanisms. We conducted extensive experiments on the COCO and OpenImages datasets, and the results demonstrate that our proposed method achieves state-of-the-art performance in layout and appearance guided generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_15252 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | LocRef-Diffusion:Tuning-Free Layout and Appearance-Guided Generation Deng, Fan Wu, Yaguang Yu, Xinyang Huang, Xiangjun Yang, Jian Yan, Guangyu Xu, Qiang Computer Vision and Pattern Recognition Artificial Intelligence Recently, text-to-image models based on diffusion have achieved remarkable success in generating high-quality images. However, the challenge of personalized, controllable generation of instances within these images remains an area in need of further development. In this paper, we present LocRef-Diffusion, a novel, tuning-free model capable of personalized customization of multiple instances' appearance and position within an image. To enhance the precision of instance placement, we introduce a Layout-net, which controls instance generation locations by leveraging both explicit instance layout information and an instance region cross-attention module. To improve the appearance fidelity to reference images, we employ an appearance-net that extracts instance appearance features and integrates them into the diffusion model through cross-attention mechanisms. We conducted extensive experiments on the COCO and OpenImages datasets, and the results demonstrate that our proposed method achieves state-of-the-art performance in layout and appearance guided generation. |
| title | LocRef-Diffusion:Tuning-Free Layout and Appearance-Guided Generation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2411.15252 |