LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zheng, Guangcong, Zhou, Xianpan, Li, Xuewei, Qi, Zhongang, Shan, Ying, Li, Xi
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909134279933952
author Zheng, Guangcong
Zhou, Xianpan
Li, Xuewei
Qi, Zhongang
Shan, Ying
Li, Xi
author_facet Zheng, Guangcong
Zhou, Xianpan
Li, Xuewei
Qi, Zhongang
Shan, Ying
Li, Xi
contents Recently, diffusion models have achieved great success in image synthesis. However, when it comes to the layout-to-image generation where an image often has a complex scene of multiple objects, how to make strong control over both the global layout map and each detailed object remains a challenging task. In this paper, we propose a diffusion model named LayoutDiffusion that can obtain higher generation quality and greater controllability than the previous works. To overcome the difficult multimodal fusion of image and layout, we propose to construct a structural image patch with region information and transform the patched image into a special layout to fuse with the normal layout in a unified form. Moreover, Layout Fusion Module (LFM) and Object-aware Cross Attention (OaCA) are proposed to model the relationship among multiple objects and designed to be object-aware and position-sensitive, allowing for precisely controlling the spatial related information. Extensive experiments show that our LayoutDiffusion outperforms the previous SOTA methods on FID, CAS by relatively 46.35%, 26.70% on COCO-stuff and 44.29%, 41.82% on VG. Code is available at https://github.com/ZGCTroy/LayoutDiffusion.
format Preprint
id arxiv_https___arxiv_org_abs_2303_17189
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation
Zheng, Guangcong
Zhou, Xianpan
Li, Xuewei
Qi, Zhongang
Shan, Ying
Li, Xi
Computer Vision and Pattern Recognition
Recently, diffusion models have achieved great success in image synthesis. However, when it comes to the layout-to-image generation where an image often has a complex scene of multiple objects, how to make strong control over both the global layout map and each detailed object remains a challenging task. In this paper, we propose a diffusion model named LayoutDiffusion that can obtain higher generation quality and greater controllability than the previous works. To overcome the difficult multimodal fusion of image and layout, we propose to construct a structural image patch with region information and transform the patched image into a special layout to fuse with the normal layout in a unified form. Moreover, Layout Fusion Module (LFM) and Object-aware Cross Attention (OaCA) are proposed to model the relationship among multiple objects and designed to be object-aware and position-sensitive, allowing for precisely controlling the spatial related information. Extensive experiments show that our LayoutDiffusion outperforms the previous SOTA methods on FID, CAS by relatively 46.35%, 26.70% on COCO-stuff and 44.29%, 41.82% on VG. Code is available at https://github.com/ZGCTroy/LayoutDiffusion.
title LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2303.17189